Abstract
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition characterized by impairments in communication, social interaction and behavior. ASD individuals develop symptoms such as recurrent actions, atypical facial expressions and challenges in social engagement. This study proposes a Multi-Atlas Context-Aware Fusion Network (MACAFNet) for ASD classification using functional connectivity (FC) features derived from resting-state functional MRI (rs-fMRI) to enable reliable ASD classification. This framework integrates information from multiple brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT to capture complementary neurofunctional features across diverse brain parcellations. Each atlas-specific connectivity representation is projected into a shared embedding space and fused using a Transformer-based attention mechanism that explicitly models inter-atlas contextual dependencies. Unlike traditional static fusion systems, this allows for dynamic and adaptive feature integration. The fused representation is then processed by a neural classifier for ASD classification. Experiments conducted on the ABIDE-I dataset demonstrate that the proposed approach achieves an accuracy of 88.46% and an AUC of 0.9546, indicating strong discriminative capability. The results highlight the effectiveness of context-aware multi-atlas fusion in capturing complex brain connectivity patterns for research-oriented ASD classification.
Keywords: Autism spectrum disorder (ASD), Resting-state fMRI (rs-fMRI), Functional connectivity (FC), Multi-atlas fusion, Transformer attention, Cross-atlas learning
Subject terms: Computational biology and bioinformatics, Diseases, Health care, Mathematics and computing, Neuroscience
Introduction
ASD is a complex neurodevelopmental syndrome characterized by difficulties in social interaction and communication flexibility. The Diagnostic and Statistical Manual of Mental Disorders has categorized ASD into three levels in reference to the support required for day-to-day functioning1. Progressive outcomes can be achieved using early identification of ASD-related patterns. Children getting early interventions show better behavioral and social flexibility2. Furthermore, cerebellar anomalies have also been connected to repetitive behaviors and decreased environmental exploration in children with ASD3. Thus, early identification and clinical assessment of neurological differences play a crucial role in improving treatment outcomes4. Few works highlight that the timeline for ASD assessment in children can be minimized when comprehensive pre-assessment information is available in advance. In contrast, ASD in adults requires more diagnostic interactions5. ASD classification is a multidisciplinary approach that includes the coordination of neurologists, psychiatrists, developmental pediatricians, and therapists6. This multidisciplinary approach provides a strong platform for integrating modern computational tools with classification and treatment.
Abnormal neurological or behavioral indicators are automatically identified using computational method, which help doctors customize treatment plans7. AI-based medical imaging assists neurologists in measuring changes in the structural and functional regions of the brain. This enables more reliable data-driven clinical decision-making8. MRI data is a promising modality for investigating the neurological basis of ASD. Structural MRI (sMRI) and functional MRI (fMRI) have aided in the detection of various abnormalities in ASD individuals. These abnormalities include irregular brain folding, excessive reduction of neural connections and poor connectivity between brain regions9. The examination of fMRI data highlights that ASD is associated with decreased inter-regional connectivity and higher segregation across functional networks10. Conventional models fail to capture neuroanatomical and functional anomalies that are better identified by modern approaches11. Enhanced volumetric and connectivity analyses using rs-fMRI provide greater support to clinicians in investigating brain developmental activities. Further, they help in examining the efficiency of therapeutic interventions over time12. Strong neurobiological biomarkers can be developed through the integration of AI with multimodal neuroimaging data13. Hence, neuroimaging-based approaches have become the basis for identifying brain-level signatures of ASD. Few studies focus on gamified mobile health applications to support children with disabilities, particularly in communication and self-management14. Mobile health and wearable technologies enable real-time behavioral monitoring and social skill training. Whereas, AI-based applications capture facial, gesture and speech patterns. However, limitations such as data confidentiality, the need for expert supervision and long-term assessment challenges remain15.
Video-based motion analysis using DL has become a viable technique for early ASD identification. It monitors head, trunk and hand movements to identify behavioral features16. The limitations of such models include data availability and generalization issues. The key features of ASD are atypical prosody, pragmatic language complications and speech impairments. Approaches such as Natural Language Processing (NLP) and speech analytics provide insights into these aspects; however, findings remain conflicting and require further investigation17. These techniques demonstrate the capability of AI in ASD diagnosis using facial, motion and speech-based datasets. But several challenges remain in clinical application. Variations exist across datasets obtained from different clinical settings. These may be due to differences in imaging procedures, feature extraction strategies and demographic composition. Many studies have developed Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN)-based pipelines for ASD classification. Still, most existing approaches either rely on voxel-level analysis or direct processing of connectivity matrices without explicitly modeling inter-atlas relationships. As a result, most existing approaches either rely on voxel-level analysis or direct processing of connectivity matrices without explicitly modeling inter-atlas relationships, limiting their ability to capture complementary information across diverse brain parcellations. Moreover, prior hybrid models typically operate as black-box classifiers. It offers limited interpretability on the contributing brain regions.
In recent years, rs-fMRI has become an important modality for analyzing alterations in FC associated with ASD. Instead of voxel-level analysis, the brain is commonly partitioned into regions of interest (ROIs) using atlases such as Automated Anatomical Labeling (AAL), Craddock 200 Functional (CC200), Dosenbach 160 Functional Regions of Interest (Dosenbach160), Eickhoff–Zilles (EZ), Harvard–Oxford (HO) and Talairach–Tournoux (TT) atlases. These atlas-based representations reduce dimensionality and improve interpretability while capturing meaningful connectivity patterns. Prior studies have shown that ASD-related abnormalities can be effectively identified through disrupted functional connectivity among distributed brain regions18,19. In parallel, single-atlas approaches such as EAG-RS leverage explainability-guided ROI selection to identify informative brain regions20. Recent works have explored multi-atlas integration to enhance ASD classification performance. Models such as MADE-for-ASD and M3ASD demonstrate that combining multiple atlases provides complementary connectivity information and improves robustness. However, effectively capturing and integrating complementary information across atlas representations remains underexplored. This results in the requirement for innovative frameworks to effectively capture complex inter-atlas relationships.
To address these limitations, this work proposes MACAFNet, a novel Multi-Atlas Context-Aware Fusion Network for ASD classification using FC features derived from rs-fMRI. Existing multi-atlas approaches depend on static or shallow fusion strategies, whereas the proposed framework models inter-atlas contextual dependencies using a lightweight Transformer Encoder. The proposed model enables dynamic weighting and interaction modeling across atlas-specific representations. This representation allows the model to effectively capture complementary neurofunctional characteristics. By integrating six widely used brain atlases, The proposed approach combines six brain atlases and presents a unified representation encompassing anatomical, functional and cytoarchitectonic perspectives. Evaluation across 63 atlas combinations demonstrates the effectiveness of the proposed fusion strategy. It achieves a strong performance on the ABIDE-I dataset.
The main contributions of this study are outlined as follows:
A novel MACAFNet framework for ASD classification using multi-atlas FC features from rs-fMRI.
A Transformer-based attention mechanism for context-aware inter-atlas feature fusion.
A multi-atlas representation integrating six brain atlases to capture complementary neurofunctional information.
An ablation study (63 atlas combinations) validating the effectiveness of multi-atlas integration.
Competitive performance on ABIDE-I (88.46% accuracy and 0.9546 AUC).
A mathematically grounded formulation ensuring reproducibility and clarity.
The remainder of this paper is organized as follows. Section 2 reviews the literature work related to ASD prediction. Section 3 presents the architecture of the proposed model and Sect. 4 describes the experimental setup. Section 5 describes the results and Sect. 6 shows the discussion. The last section presents the conclusion.
Related work
Recent studies employ deep learning (DL) techniques on neuroimaging, behavioral and linguistic data in ASD screening research. MRI-based models aid in identifying structural and functional abnormalities associated with ASD. Eye-tracking data, facial expression analysis and motion-based features support automated screening systems. Speech and language variances are analyzed using NLP. The gamified mobile health applications increase accessibility and engagement for ASD children. Few works on ensemble and multimodal frameworks have proved enhanced predictive performance by combining diverse data sources. Despite these advancements, challenges remain in terms of dataset diversity, model generalization and clinical adoption. This section reviews various approaches for ASD detection, their strengths and limitations to provide a clear understanding of recent developments in this field. Table 1 presents an overview of models ranging from conventional machine learning (ML) approaches to advanced DL and multimodal frameworks used for ASD detection.
Table 1.
Comprehensive summary of existing studies that utilize ML, DL and multimodal data integration techniques to support ASD classification and screening.
| Author | Model/method | Data Modality | Dataset | Pros | Cons |
|---|---|---|---|---|---|
| Santana et al.12 | SVM, ANN, other classifiers (meta-analysis) | rs-fMRI | 55 rs-fMRI studies | SVM robust; ANN scales with larger datasets; multimodal approaches improved sensitivity to 84.7% | Accuracy decreased with larger samples; poor methodological quality; limited clinical applicability |
| Ning Qiang et al.21 | HRVAE + LASSO | rs-fMRI | 871 subjects from ABIDE | Captures hierarchical brain networks; 82.1% accuracy; promising for clinical assessment support | Complexity; potential overfitting; requires large datasets |
| Hazlett et al.22 | DL on cortical surface hyper-expansion | sMRI | High-risk infants | Early prediction at 24 months; 88% sensitivity; non-invasive biomarkers | Small cohort; longitudinal follow-up required |
| Bengs et al.23 | 4D spatio-temporal DL (CNN + Conv-Recurrent) | fMRI | fMRI data | Captures spatial and temporal features; F1-score 0.71 | Moderate performance; high computational cost |
| Sachdeva et al.24 | MLP on connectivity + graph topology features | rs-fMRI + T1-weighted MRI | rs-fMRI + T1-weighted images | 83.57% accuracy; 0.978 AUC; identifies key ROIs | Feature engineering needed; limited interpretability |
| Heinsfeld et al.25 | DNN on functional connectivity matrices | rs-fMRI | 505 ASD individuals and 530 TC from ABIDE I | One of the earliest DL models for ASD prediction using ABIDE; achieved 70% accuracy; demonstrated feasibility of DL on connectivity features | Performance limited; sensitive to site variability |
| Parisot et al.26 | Graph Convolutional Network (GCN) using population graph | rs-fMRI | 871 subjects from ABIDE, 731 subjects from ADNI | Demonstrates the relationships between subjects; improves ASD classification; integrates imaging and non-imaging data | Requires graph construction; performance depends on population graph design |
| Liu et al.27 | A Stacked Sparse Denoising Autoencoder (SSDAE) and Multi-Layer Perceptron (MLP) |
rs-fMRI (Multi-atlas) |
505 ASD individuals and 530 TC from ABIDE I | Higher accuracy than prior ABIDE studies; strong ASD detection; excellent performance on curated subset | Lower specificity; generalization variability; subset bias risk; requires external validation |
| Yang et al.28 | Multi-Atlas Multi-View Learning | rs-fMRI | 407 ASD and 468 TC, 6 to 64 years | Integrates connectivity patterns from multiple atlases; improves representation learning | Limited modeling of inter-atlas contextual relationships |
| Jung et al.29 | Explainability-guided ROI selection (EAG-RS) framework | rs-fMRI | 539 ASD and 573 TD, aged 7 to 64 years. | Explainability-guided ROI selection; captures high-order inter-regional relationships; provides neuroscientific insights into ASD subtypes | MLP-based architecture with high parameter complexity; relies only on FC (single modality); limited multimodal integration. |
| Federica Cilia et al.30 | CNN on scanpath visual maps | Eye-tracking (gaze scan-paths) | 59 school-age participants | Non-invasive; targets social attention; 90% accuracy | Small sample; variability in stimuli; limited external validity |
| Kanhirakadavath et al.31 | DNN, SVM, RF on ETSP images | Eye-tracking spatial plots (ETSP) | Child eye-tracking dataset | DNN achieved 97% AUC (augmented); captures gaze dynamics; non-invasive | Small cohort; task/stimuli dependent; dataset variability |
| Farhat et al.32 | Ensemble CNN (VGG16 + Xception) | Facial images | Kaggle ASD Face Image | High accuracy (97–99%); strong generalization | Limited dataset diversity; may not generalize clinically |
| Anderson et al.33 | Markerless video-based gait analysis (OpenPose) | Video (gait and movement) | Children 1–5 years with/without NDD | Non-invasive; accurate gait metrics; early-age feasibility | Step length affected by body size; methodological factors |
| Jin et al.34 | ML on home-based video | Video (home recordings) | 9,959 subjects; 19 studies | High sensitivity/specificity; user-friendly; early ASD detection | Variable video quality; feature selection may affect results |
| Dia et al.35 | Transformer-based affect recognition | Video (facial and motion cues) | SSBD dataset; YouTube videos | Continuous affect recognition; ASD-specific dataset; highlights stimming & face analysis | Data quality/diversity issues; model underperforms if trained on neurotypical data only |
| Ganjigunte et al.36 | BAR pipeline | Video (behavioral actions) | 400 ASD + 125 ODD children; SSBD | Untrimmed video analysis; 78–81% accuracy; clinically relevant actions | Requires labeled actions; dataset-specific; computationally intensive |
| Abbas et al.37 | Multi-modular ML assessment (parent + video + clinician) | Multimodal (video, clinical, parental) | 375 children, 18 to 72 months | Outperforms conventional screeners; AUC improvement; time-efficient | Needs validation in general populations; limited primary care testing |
| Ma et al.38 | ML on natural speech prosody | Speech/audio | Meta-analysis | Pitch-related features distinguish ASD; 75–80% sensitivity/specificity | Temporal features less informative; age-dependent effects |
| Chi et al.39 | RF and CNN | Speech/audio | Self-recorded child speech (mobile game) | CNN achieved 79% accuracy; works in home environment; scalable | Audio quality variable; requires preprocessing |
| Mahmoudi et al.14 | Scoping review on gamified apps | Mobile health (mHealth) applications | 38 studies covering 32 apps | Increases engagement/motivation; supports communication & education | Mixed effectiveness; some disability groups underexplored; long-term impact unknown |
The reviewed literature works demonstrate ML and DL approaches for ASD detection across multiple modalities. Studies based on rs-fMRI neuroimaging have shown promising results in identifying FC alterations and reliable biomarkers. Furthermore, non-invasive approaches such as eye-tracking, facial analysis, video-based behavior recognition and speech processing provide scalable solutions for early screening. Many neuroimaging studies rely on single-atlas representations or simple connectivity features. This limits the ability of the models to capture the complex and heterogeneous nature of brain connectivity. Recent multi-atlas and explainability-guided methods also do not fully exploit contextual relationships across atlas representations. Also, the challenges such as dataset heterogeneity and limited generalization remain significant concerns.
Methods
The methods section describes the proposed Multi-Atlas Context-Aware Fusion Network (MACAFNet) developed for ASD classification using rs-fMRI connectivity features. The framework utilizes FC representations derived from multiple brain atlases to capture complementary neurofunctional information across different brain parcellations. This context-aware fusion strategy allows the model to effectively capture inter-atlas dependencies and enhances predictive robustness and generalization for rs-fMRI-based ASD classification.
Proposed architecture
The proposed architecture MACAFNet, is designed to classify ASD-related patterns using FC features derived from rs-fMRI data. The framework utilizes atlas-based brain parcellations to represent neural interactions among predefined ROIs, instead of directly processing MRI slices. Six brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT are employed to generate complementary FC representations of the brain. In this study, each subject’s ROI time series is normalized using z-score normalization. It is then followed by FC computation using Pearson correlation. The upper triangular elements of the FC matrix are extracted as feature vectors. These connectivity features are stored and used as input vectors to the model. These features are transformed into a shared embedding space. This enables consistent representation across different atlas configurations. A Transformer-based attention module is employed to model contextual relationships among the atlas representations and learn the relative importance of each atlas in the classification process.
The attention mechanism facilitates context-aware fusion of multi-atlas connectivity features. This allows the network to capture complementary information from different brain parcellation schemes. Each atlas-specific FC vector is independently projected into a shared embedding space using fully connected layers. The resultant embeddings are stacked and passed through a Transformer encoder to model inter-atlas dependencies. A lightweight Transformer configuration with one layer, four heads and embedding size of 64 is adopted to balance model complexity and generalization. This is configured to adapt to the limited sample size of neuroimaging datasets. Finally, mean pooling is applied across atlas embeddings to obtain a unified representation for classification. This architecture enables the model to effectively capture complex inter-regional connectivity patterns and improves the robustness and generalization capability of rs-fMRI-based ASD prediction. Figure 1 illustrates the overall workflow of the proposed MACAFNet framework for ASD classification.
Fig. 1.
Architecture of the proposed framework used for ASD classification. Multi-atlas rs-fMRI connectivity features are embedded into a shared space and fused using a Transformer-based attention module and fully connected layers for robust ASD classification.
Dataset description and experimental data splitting strategy
The proposed framework is evaluated on the publicly available Autism Brain Imaging Data Exchange I (ABIDE-I) dataset40. It consists of rs-fMRI data collected from 17 different international sites. The sites are NYU Langone Medical Center (NYU), University of Michigan (UM), University of Utah School of Medicine (USM), University of California Los Angeles (UCLA), Kennedy Krieger Institute (KKI), Oregon Health & Science University (OHSU), Stanford University (SU), University of Pittsburgh School of Medicine (PITT), Yale Child Study Center (YALE), Trinity College Dublin (TRINITY), University of Leuven (LEUVEN), University of Cambridge (UC), California Institute of Technology (CALTECH), Social Brain Lab BCN NIC UMC Groningen (SBL), San Diego State University (SDSU), University of Miami (UM), and Max Planck Institute for Human Cognitive and Brain Sciences (MAX_MUN). The original dataset contains 1112 subjects. This work employed 1035 subjects with complete phenotypic and imaging information. The final dataset consists of 505 ASD and 530 typically developing (TD) controls. The preprocessed functional data were obtained from the Preprocessed Connectomes Project (PCP) repository41,42, specifically using derivatives generated through the Configurable Pipeline for the Analysis of Connectomes (CPAC) with the filt_global strategy (global signal regression). ROI-based time-series data corresponding to multiple atlases (CC200, AAL, Dosenbach160, EZ, TT, and HO) were directly utilized from the provided .1D derivative files.
Pairwise Pearson correlation was applied to standardized ROI time-series signals to construct functional connectivity (FC) matrices. The upper triangular elements of these matrices were extracted and vectorized to form atlas-specific FC feature vectors. These atlases provide complementary perspectives of brain organization such as anatomical, functional and cytoarchitectonic representations. This helps in enabling a richer characterization of brain connectivity patterns. Diagnostic labels were obtained from the phenotypic metadata. It consists of a column named “DX_GROUP” which defines the ground truth (1: ASD, 2: TD). These labels were converted into a binary format, with ASD as ‘0’ and TD as ‘1’ for carrying out the training process. These labels were associated with the corresponding connectivity features to construct input–target pairs. Subjects with missing diagnostic labels were excluded from the analysis.
Let
denote the number of ROIs in atlas a. The FC matrix for each atlas is represented as a symmetric matrix
, where each element corresponds to the Pearson correlation coefficient between ROI time series. To eliminate redundancy, only the upper triangular elements of the FC matrix are retained. This resulted in a feature vector of size
for each atlas. For each atlas combination experiment, the dataset is randomly partitioned into training (70%), validation (15%), and testing (15%) subsets using a stratified splitting strategy to preserve class distribution, thereby preventing site-specific data leakage. A fixed random seed (42) is used to ensure reproducibility across all experiments.
Multi-atlas feature representation
The study incorporates multiple brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT to capture complementary aspects of brain organization such as anatomical structure, functional connectivity and cytoarchitectonic properties. Each atlas provides a unique perspective on brain parcellation. It helps in the extraction of diverse and informative features. This multi-atlas strategy enhances the representation of complex neural patterns and improves the ability to detect subtle abnormalities associated with ASD. It leads to more robust and reliable classification performance compared to conventional single-atlas approaches. The number of ROIs and corresponding feature dimensions vary across atlases. This enables the model to capture brain connectivity patterns at different spatial resolutions as outlined in Table 2.
Table 2.
Summary of the brain atlases used in this study.
| Atlas | ROIs | Feature dimension | Representation type | Regions captured | Relevance to ASD analysis |
|---|---|---|---|---|---|
| AAL | 116 | 6670 | Anatomical | Divides the brain into anatomically defined regions based on structural boundaries | Helps analyze structural-functional disruptions commonly observed in ASD |
| CC200 | 200 | 19900 | Functional | Clusters brain regions based on functional similarity from rs-fMRI signals | Captures fine-grained functional connectivity patterns altered in ASD |
| Dosenbach160 | 161 | 12880 | Functional | Defines ROIs based on task-control and functional network regions | Useful for studying attention, control networks, and cognitive impairments in ASD |
| EZ | 116 | 6670 | Cytoarchitectonic | Based on cellular structure and cortical microarchitecture | Enables analysis at a finer biological level relevant to neurodevelopmental disorders |
| HO | 111 | 6105 | Anatomical | Probabilistic atlas derived from structural MRI data | Provides robust anatomical reference for cross-subject consistency |
| TT | 97 | 4656 | Anatomical | Classical coordinate-based brain atlas using standardized space | Facilitates spatial normalization and comparison across studies |
It outlines the number of ROIs, feature dimensionality, representation type, the type of brain regions captured by each atlas and their relevance to ASD analysis.
Each atlas captures brain organization at a different spatial scale. Thus, the integration of multi-atlases enables multi-scale feature representation learning. This method combines fine-grained local connectivity patterns with coarse-grained global brain network structures. The atlas-specific FC feature vectors are independently projected into a shared embedding space and subsequently fused using a Transformer-based attention mechanism. This results in improved discriminative capability for ASD classification.
Functional connectivity (FC) extraction
Functional connectivity represents statistical relationships between spatially separated brain regions. This is extensively used to depict neural communication patterns in rs-fMRI data. In this study, FC vectors are computed from ROI time series obtained through the PCP41 using Pearson correlation. For each subject, the brain is parcellated into ROIs using six brain atlases such as AAL, CC200, Dosenbach160, EZ, HO and TT. Each atlas captures the neural activity at different levels of granularity. These atlases provide a unique spatial partitioning of the brain. The corresponding FC representations are derived from ROI-level time-series using standard preprocessing pipelines. These are provided as vectorized connectivity features. For each atlas, a symmetric FC matrix is constructed. But only the upper triangular elements are extracted to form the feature vector. This eliminates the redundant information. These connectivity features represent pairwise statistical dependencies between brain regions and capture large-scale brain network interactions associated with ASD. The resulting connectivity feature vectors vary in dimensionality as each atlas defines a different number of ROIs. These atlas-specific FC feature vectors are independently used as inputs to the embedding layers of the proposed MACAFNet framework.
Multi-atlas context-aware fusion network (MACAFNet) module
MACAFNet is designed to combine connectivity representations obtained from multiple brain atlases to capture complementary functional patterns associated with ASD. The atlas-specific connectivity features are transformed into compact embeddings using linear projection layers with ReLU activation and dropout regularization. These embeddings capture high-level representations of FC patterns while reducing redundancy and computational complexity. The embeddings from all atlases are then stacked and processed using a Transformer-based attention mechanism.
This module enables the model to capture relationships between different brain atlases by leveraging an attention-based weighting mechanism. It adaptively emphasizes informative connectivity patterns while suppressing less relevant features. This process helps in enhancing the feature representation. A Transformer encoder is employed to model inter-atlas dependencies it consists of a single layer with four attention heads, an embedding dimension of 64 and a dropout rate of 0.4. This configuration provides a balance between model capacity and generalization. The outputs from the Transformer are subsequently aggregated using mean pooling to obtain a unified feature representation for final classification. This final representation consists of individual atlas information and their relationships. This gives a clear understanding of brain connectivity. A dropout layer is employed to reduce overfitting. ReLU activation is used in the embedding layers to help the model learn more complex relationships from the input features. Figure 2 outlines the workflow of MACAFNet fusion.
Fig. 2.
MACAFNet architecture showing atlas-wise embedding, Transformer-based cross-atlas attention, feature fusion via pooling, and fully connected layers for ASD classification.
Classification and model training
Once the multi-atlas feature fusion is completed, the integrated representation is passed through fully connected layers to classify each subject into ASD or typically developing (TD) categories for research analysis. Each fused feature vector is processed through a ReLU-activated dense layer. It is followed by a dropout layer with a rate of 0.4 to reduce overfitting and improve generalization. The final output layer produces logits for ASD classification. The model is optimized using binary cross-entropy with logits loss (BCEWithLogitsLoss) with class weighting to address class imbalance. It ensures numerical stability during training. To address class imbalance in the dataset, a class weighting strategy is applied based on the distribution of ASD and TD samples in the training set. Adam optimizer is employed to provide efficient adaptive learning rate updates for stable convergence. The initial learning rate is set to 0.001.
Training is performed using mini-batch gradient descent with a batch size of 32. This batch size is chosen to balance computational efficiency and stable gradient estimation under limited dataset conditions. The model is trained for a maximum of 50 epochs to ensure sufficient convergence. To prevent overfitting, early stopping with a patience of 7 epochs is employed. Moreover, a ReduceLROnPlateau scheduler reduces the learning rate by a factor of 0.5 when validation performance does not improve for three consecutive epochs. A stratified train–validation–test split is performed while ensuring that subjects from the same acquisition site are not distributed across different splits, thereby preventing site-specific data leakage. The model achieving the highest validation F1-score is selected as the best model. The held-out test set is used only for final unbiased evaluation. This is done to ensure no data leakage.
For each experiment, the dataset is split into training (70%), validation (15%), and testing (15%) subsets using a stratified sampling strategy to preserve class distribution. Experiments are conducted across all 63 possible combinations of the six brain atlases, ranging from single-atlas models to full multi-atlas fusion. During training, the network learns discriminative functional connectivity (FC) patterns from each atlas while MACAFNet explicitly models inter-atlas contextual dependencies using a lightweight Transformer encoder, enabling dynamic weighting of atlas-specific connectivity patterns rather than static fusion. This integrated design enables MACAFNet to capture complex brain connectivity patterns. This results in an enhanced ASD classification performance. To address class imbalance and enhance classification robustness, the decision threshold is not fixed at 0.5. Instead, the optimal threshold is determined by maximizing the F1-score on the validation set. During inference, the output logits are passed through a sigmoid activation function. The optimized threshold is applied to obtain final ASD/TD predictions on the test set. All experiments are conducted with a fixed random seed of 42 to ensure reproducibility of results.
Mathematical modeling of the proposed architecture
The overall pipeline is mathematically formulated to systematically transform raw rs-fMRI signals into discriminative multi-atlas representations, followed by attention-based fusion using the proposed MACAFNet architecture. Let the dataset consist of N subjects. For each subject, functional connectivity matrices are extracted from six brain atlases.
Data preparation
Raw rs-fMRI signals are preprocessed and organized into labeled samples. Each subject is associated with multiple atlas-based representations. Preprocessing includes denoising, filtering and motion correction. The complete dataset D after preprocessing and organizing is given in Eq. (1).
![]() |
1 |
where,
represents total number of subjects,
is rs-fMRI data of subject s and
denotes the label (0 = Control, 1 = ASD).
Dataset splitting
The dataset is evaluated using stratified split to ensure reliable and unbiased performance estimation as represented in Eq. (2). This split helps to preserve class distribution across all subsets. Further, it ensures balanced and unbiased model training and evaluation.
![]() |
2 |
ROI time-series extraction
For each subject and atlas, brain activity is represented as time-series signals extracted from predefined ROIs, capturing temporal neural dynamics. Brain signals are extracted from regions of interest across time points using Eq. (3).
![]() |
3 |
where,
denotes ROI time-series matrix for subject s,
represents the number of ROIs (depending on the atlas) and T denotes the number of time frames.
Feature vector construction
The upper triangular elements of the symmetric connectivity matrix are extracted and vectorized to form a compact feature representation by removing redundant and self-correlated values, which is denoted in Eq. (4).
![]() |
4 |
where,
denotes the feature vector for subject s,
is the extraction of upper triangle and
represents the vectorization operator.
Feature normalization
Feature normalization helps to ensure consistent feature scaling across subjects. It is represented in Eq. (5).
![]() |
5 |
Multi-atlas feature representation
Multiple atlas-specific feature vectors are generated for each subject. Features from multiple atlases are combined into a unified representation as represented in Eq. (6).
![]() |
6 |
where, M represents the number of atlases,
is the feature vector from atlas ‘i’, i ∈ {1, 2, …, M} are the atlas index and
is the fused feature vector.
Feature embedding
Each atlas-specific feature vector is transformed into a common latent space using a linear projection followed by a non-linear activation function. This ensures that features from different atlases have a consistent representation. This enables effective learning of inter-atlas relationships. Equation (7) represents the embedded feature.
![]() |
7 |
where,
represents the embedded feature,
is the non-linear Rectified Linear Unit activation function,
is the learnable weight matrix,
represents the input feature vector from atlas ‘i’ and
is the bias term.
Sequence formation
The embedded feature vectors from all atlases are stacked to form an ordered sequence, which serves as input to the attention mechanism. This sequence representation given in Eq. (8) allows the model to learn relationships and dependencies across different atlases.
![]() |
8 |
where,
is the embedding matrix and
is the embedded vector.
Query, key, value projection
The embedding matrix is linearly transformed into Query, Key, and Value representations using learnable weight matrices. These projections enable the model to compute attention scores and capture relationships between atlas features. The formula for transformation is given in Eqs. (9), (10) and (11) respectively.
![]() |
9 |
![]() |
10 |
![]() |
11 |
where,
denotes the query matrix (represents “what to attend”),
denotes the key matrix (represents “what is available”),
denotes the value matrix (represents “actual information”),
denotes the input embedding matrix (stacked atlas features),
are the learnable weight matrix for Query, Key and Values respectively,
Attention computation
Attention weights are computed by measuring the similarity between Query and Key representations. It is followed by normalization using the SoftMax function. These weights determine the relative importance of each atlas feature. These computed weights are then applied to the Value representations to generate context-aware features that capture inter-atlas relationships. It is given in Eqs. (12) and (13).
![]() |
12 |
![]() |
13 |
where,
represents the attention weight matrix,
, Similarity matrix between Query and Key,
denotes the scaling factor to stabilize gradients,
: Attention output (context-aware representation) and
denotes the value matrix.
Multi-head attention
Multi-head attention applies multiple parallel attention mechanisms to the same input. It allows the model to learn diverse and complementary inter-atlas relationships. Each attention head focuses on different aspects of the feature interactions. Finally, their outputs are combined to form a richer and more informative representation as given by Eq. (14).
![]() |
14 |
where,
denotes the final multi-head attention output,
is the output of the attention heads,
is the output projection matrix, H is the number of attention heads and
is the Concatenated output.
Context-aware fusion
The attention outputs from all atlases are aggregated to form a single unified feature vector that captures both individual atlas characteristics and their interrelationships. This fusion step represented in Eq. (15) integrates context-aware information into a compact representation for effective classification.
![]() |
15 |
where,
: Final fused feature vector.
Classification layer
The fused feature vector
is transformed through a fully connected layer with ReLU activation to obtain a high-level hidden representation
, which is given in Eq. (16).
![]() |
16 |
where,
is the weight matrix,
is Bias vector of the first fully connected layer.
The hidden representation
is projected through a linear layer to produce logits
for binary ASD classification, which is given in Eq. (17).
![]() |
17 |
where,
is the weight matrix of the output layer,
is the bias vector of the output layer.
Loss function
The model is optimized using a weighted binary cross-entropy with logits loss, which combines sigmoid activation and cross-entropy in a numerically stable formulation. Class weights are incorporated to address imbalance between ASD and TD samples, ensuring balanced learning and improved sensitivity to minority class predictions. The loss function
is represented in Eq. (18).
![]() |
18 |
where,
is the contribution when true label is ASD,
is the contribution when true label is Control and
represents the positive class weight.
Optimization process (Adam optimizer)
Adaptive moment estimation (Adam) is an efficient optimization procedure. It combines momentum and adaptive learning rates. Momentum supports smooth gradients via a moving average. Adaptive learning rates adjust the learning rate for each parameter on the basis of its history. The first momentum called the mean
, defines the exponential moving average of the gradients. It reduces noise updates and allows faster convergence. It is given by Eq. (19).
![]() |
19 |
where
refers to the decay rate (i.e., 0.9 for the momentum term),
is the exponential moving average of past momentum, and
refers to the gradient of the loss at time step t.
The second moment called the variance
tracks of the magnitude of the gradient over time. It is shown in Eq. (20).
![]() |
20 |
where
represents the decay rate (i.e.) 0.999 for variance term and
is the exponential moving average of the squared gradients.
The bias-corrected first moment
helps in removing the initialization bias in the early stage of the training steps. It is initialized to zero when the training starts. The correction bias is adjusted via
. It is given in Eq. (21).
![]() |
21 |
The bias-corrected second moment
helps in removing the initial bias toward zero.
helps to correct bias which is defined via Eq. (22).
![]() |
22 |
With respect to the parameter update
, the model weights are updated. It integrates the momentum and adaptive learning rate to produce a balanced and fast converging optimization step as defined in Eq. (23).
![]() |
23 |
![]() |
Model selection
The optimal model is selected based on best F1-score validation performance. It is denoted in Eq. (24) is identified by maximizing the validation F1-score, ensuring a balanced trade-off between precision and recall for ASD classification.
![]() |
24 |
Proposed methodology
This section presents the proposed methodology for ASD classification using multi-atlas FC features derived from rs-fMRI data. The framework consists of a structured pipeline where FC matrices are extracted from six different brain atlases. It is followed by atlas-specific feature embedding, contextual feature learning through a Transformer encoder, multi-atlas fusion and final classification. Brain atlases were employed to obtain complementary FC representations of brain networks. For each atlas, ROI-wise time-series signals were extracted and pairwise Pearson correlation coefficients were computed to generate connectivity matrices. These matrices were transformed into feature vectors by extracting the upper triangular elements of the symmetric FC matrices to remove redundancy and projected into a shared embedding space. The proposed MACAFNet framework integrates atlas-specific features using a Transformer-based attention mechanism with a single encoder layer and four attention heads that captures contextual dependencies across different atlas representations. This enables the model to learn complementary connectivity patterns associated with ASD. The fused representation is then passed through a fully connected classification layer to predict whether a subject belongs to the ASD or typically developing class. The final output consists of logits, which are converted to probabilities using a sigmoid activation function during inference. The detailed workflow of the proposed approach is summarized in Algorithm 1.
Algorithm 1.
MACAFNet multi-atlas ASD classification model
Figure generation and visualization tools
All figures in this study were generated using a combination of tools. The architectural diagrams were created using draw.io (diagrams.net) for clear schematic representation of the proposed framework. Performance plots, including accuracy curves, loss curves, and ROC curves, were generated using Python-based visualization libraries, specifically Matplotlib and Scikit-learn visualization utilities. Data processing and visualization were performed in a Python 3.8 environment.
Results
This section presents the experimental evaluation of the proposed MACAFNet for ASD classification using multi-atlas FC features derived from rs-fMRI data. All experiments were conducted on an NVIDIA GPU server (NVIDIA-SMI 535.183.01) equipped with four NVIDIA RTX A6000 GPUs (48 GB VRAM each), running CUDA 12.2.
Comparison with existing studies
The performance of the proposed MACAFNet framework was compared with several prior studies that utilized ABIDE I dataset consisting of 1,035 subjects for ASD classification. Table 3 outlines the performance measure of the existing models and the proposed MACAFNet framework in terms of test accuracy. AE + DNN25, CNN-based model43 and AE + SLP44 produced an accuracy of 70% approximately, while SAE + MLP45 produced 70.8% and MADE-for-ASD27 produced 75.2%. The proposed MACAFNet achieves an accuracy of 88.46% for the atlas combination CC200 + EZ + TT. This improvement demonstrates the effectiveness of leveraging complementary connectivity information from multiple brain parcellations. It also emphasizes the benefit of context-aware multi-atlas fusion in capturing discriminative neural patterns associated with ASD. Furthermore, MACAFNet dynamically models inter-atlas dependencies leading to improved generalization. The results clearly indicate that MACAFNet outperforms existing approaches by providing a more reliable and robust framework for ASD classification model. The reported performance corresponds to the best-performing model selected based on the highest validation F1-score (0.9252) during training with early stopping, and evaluated on a held-out test set.
Table 3.
Comparation of performance of the proposed MACAFNet framework with the existing ASD classification methods on the ABIDE I CPAC dataset.
| Study | Model | Number of subjects | Accuracy (%) |
|---|---|---|---|
| Heinfield et al. 25 | AE + DNN | 1035 | 70 |
| Sherkatghanad et al.43 | CNN | 1035 | 70.22 |
| Eslami et al.44 | AE + SLP | 1035 | 70.3 |
| Almuqhim et al.45 | SAE + MLP | 1035 | 70.8 |
| Liu et al. 27 | MADE-for-ASD | 1035 | 75.2 |
| Proposed MACAFNet | MACAFNet + Multi-Atlas Context-Aware Fusion | 1035 | 88.46 |
Ablation study: effect of multi-atlas combinations
An ablation study was conducted using all possible 63 combinations of six brain atlases to evaluate the contribution of individual atlases and their combinations. The combinations range from single-atlas models to multi-atlas fusion. The results demonstrate that multi-atlas fusion consistently outperforms single-atlas models across all metrics. Among all combinations, CC200 + EZ + TT achieved the highest performance, with a test accuracy of 88.46%, F1-score of 0.8448, and AUC of 0.9546, while maintaining balanced precision (0.8448) and recall (0.8448). High sensitivity (0.8448) and specificity (0.9082) indicate effective detection of both ASD and TD subjects. The performance gain is not strictly monotonic with the number of atlases, indicating that the quality and complementarity of selected atlases are more important than simply increasing the number of inputs. The performance metrics of the top 10 multi-atlas combinations are shown in Table 4. All results reported in this study correspond to models trained using a stratified 70%–15%–15% split (training–validation–test), with class distribution preserved across splits. The best model for each atlas combination is selected based on validation F1-score and evaluated on the held-out test set to ensure fair and unbiased comparison.
Table 4.
Performance metrics of the top 10 multi-atlas combinations used for ablation study.
| Atlas combinations | Test accuracy | Training accuracy | Validation accuracy | Precision | F1-score | AUC | Sensitivity | Specificity |
|---|---|---|---|---|---|---|---|---|
| CC200 + EZ + TT | 0.8846 | 0.9876 | 0.9290 | 0.8448 | 0.8448 | 0.9546 | 0.8448 | 0.9082 |
| AAL + CC200 + HO | 0.8782 | 0.9848 | 0.9097 | 0.8679 | 0.8288 | 0.9229 | 0.7931 | 0.9286 |
| AAL + CC200 + Dosenbach160 + HO + TT | 0.8718 | 0.9834 | 0.9226 | 0.7879 | 0.8387 | 0.9372 | 0.8466 | 0.8571 |
| AAL + CC200 + Dosenbach160 + EZ + TT | 0.8654 | 0.9820 | 0.9032 | 0.7937 | 0.8264 | 0.9367 | 0.8621 | 0.8673 |
| AAL + CC200 | 0.8590 | 0.9751 | 0.9032 | 0.8000 | 0.8136 | 0.9263 | 0.8276 | 0.8776 |
| AAL + CC200 + EZ + HO + TT | 0.8526 | 0.9738 | 0.9161 | 0.7869 | 0.8067 | 0.9340 | 0.8276 | 0.8673 |
| AAL + Dosenbach160 + EZ | 0.8462 | 0.9682 | 0.9484 | 0.7361 | 0.8154 | 0.9305 | 0.8138 | 0.8061 |
| AAL + EZ + HO | 0.8397 | 0.9862 | 0.9032 | 0.7143 | 0.8148 | 0.9335 | 0.8483 | 0.7755 |
| Dosenbach160 + HO | 0.8269 | 0.9751 | 0.9161 | 0.7123 | 0.7939 | 0.9353 | 0.8448 | 0.7857 |
Significant values are in bold.
The ablation study highlights several key findings:
Single atlases such as CC200 provide strong baseline performance (accuracy 87.18%, F1-score 0.8438), confirming their informative connectivity representation.
Adding complementary atlases improves recall and sensitivity, showing enhanced detection of ASD subjects.
Precision-recall trade-offs are observed across different atlas combinations, emphasizing the value of context-aware fusion to balance false positives and false negatives.
The proposed attention-based fusion mechanism effectively learns optimal weighting across atlas features, avoiding redundancy and enhancing discriminative feature learning.
Additionally, to handle class imbalance and improve robustness, the classification threshold is optimized by maximizing the F1-score on predicted probabilities rather than using a fixed threshold of 0.5. This strategy contributes to improved balance between sensitivity and specificity across atlas combinations.
The training and validation accuracy and loss curves for the CC200 + EZ + TT atlas combination outlined in Fig. 3 demonstrates smooth convergence behavior. A strong generalization capability is obtained with minimal divergence between training and validation accuracy. The training loss and validation loss gradually decreased highlighting effective optimization, model convergence, reduced overfitting and improved generalization. The ROC curve depicted in Fig. 4 shows an AUC of 0.9546, confirming excellent class separability. This high AUC indicates strong discriminative ability and robust ranking performance across varying decision thresholds.
Fig. 3.
Training and validation accuracy (a) and loss (b) curves of the MACAFNet model for the CC200 + EZ + TT atlas combination. The curves show smooth convergence and minimal overfitting.
Fig. 4.

ROC curve of the MACAFNet model (CC200 + EZ + TT), with AUC = 0.9546, demonstrating strong discriminative ability for ASD classification.
Hyperparameter tuning performed using various ranges of batch size and epoch revealed that there was a slight decrease in the performance of the model for different batch sizes of {4, 8, 16, 64} and epochs between (40, 100). Specifically, smaller batch sizes introduced noisy gradient updates, while larger batch sizes and increased epochs led to overfitting. The model produced peak performance at a batch size of 32 and epoch of 50. This indicates an optimal balance between convergence stability and generalization.
Error analysis and model performance
The performance of the top MACAFNet model (CC200 + EZ + TT) was further analyzed using the confusion matrix to understand misclassifications between ASD and TD subjects, which is depicted in Fig. 5. The confusion matrix corresponds to the best-performing model (selected using validation F1-score) evaluated on the held-out test set.
Fig. 5.

Confusion matrix of the MACAFNet model (CC200 + EZ + TT), showing the distribution of true and predicted classes for ASD and TD subjects.
The performance of the top MACAFNet model (CC200 + EZ + TT) was further analyzed using the confusion matrix to understand misclassifications between ASD and TD subjects. From the results it is evident that the model correctly classified 89 TD and 49 ASD subjects. It misclassified 9 TD subjects as ASD and 9 ASD subjects as TD. This indicates that the model maintains a balanced performance across both classes. It effectively detects ASD subjects while minimizing false positives among TD subjects. The error analysis highlights that the proposed MACAFNet achieved high sensitivity (0.8448) and specificity (0.9082). Despite strong performance, a small number of misclassifications may arise due to inter-subject variability, site-specific differences in data acquisition and overlapping connectivity patterns between ASD and TD groups. Further, training and validation curves exhibit smooth convergence with minimal divergence, indicating stable learning and good generalization. The top-performing atlas combination achieved a test accuracy of 88.46%, F1-score of 0.8448, and AUC of 0.9546. This indicates consistent and robust predictive capability.
Discussion
The experimental results demonstrate that the proposed MACAFNet effectively leverages complementary FC information from multiple brain atlases to improve ASD classification. Multi-atlas fusion consistently outperforms single-atlas models across all metrics. The CC200 + EZ + TT combination achieved the highest accuracy of 88.46%, with a recall of 0.8448 and an AUC of 0.9546, indicating strong discriminative capability and generalization. A key observation is that different atlases capture distinct aspects of brain organization. The superior performance of CC200 + EZ + TT is due to their complementarity, where CC200 captures fine-grained functional connectivity, EZ provides cytoarchitectonic information, and TT contributes stable anatomical structure. This combination enables the model to learn multi-level representations of brain connectivity. Importantly, increasing the number of atlases does not always improve performance; instead, selecting complementary atlases is more effective. The Transformer-based attention mechanism further enhances performance by learning inter-atlas dependencies and emphasizing the most informative features while suppressing noise. The model’s robustness is supported by effective training strategies such as dropout, early stopping, and learning rate scheduling, which improve convergence and reduce overfitting. From a clinical perspective, the achieved sensitivity and specificity indicate strong potential for reliable ASD classification. Although the proposed model achieves strong performance, variability across acquisition sites in ABIDE remains a challenge. Future work will explore domain adaptation and site-invariant learning strategies.
Conclusion
This study presented a MACAFNet for ASD classification using functional connectivity features derived from rs-fMRI data. The framework effectively integrates connectivity representations from multiple brain atlases and leverages a transformer-based fusion mechanism to capture contextual dependencies among atlas-specific features. Experimental evaluation across 63 atlas combinations, using validation-guided model selection and held-out test evaluation, demonstrated that multi-atlas fusion consistently outperforms single-atlas models. Notably, the CC200 + EZ + TT combination achieved the highest performance, with an accuracy of 88.46% and an F1-score of 0.8448. The results confirm that combining complementary atlas representations enables richer and more discriminative feature learning, while the attention-based fusion mechanism enhances the modeling of inter-atlas relationships. Overall, MACAFNet provides a robust and scalable approach for ASD detection using rs-fMRI data. Experimental evaluation across 63 atlas combinations, using validation-guided model selection and held-out test evaluation, demonstrated that multi-atlas fusion consistently outperforms single-atlas models. Despite the promising results, the study is limited by dataset heterogeneity and sample size constraints inherent to ABIDE-I. Future work will focus on integrating multi-modal neuroimaging data, exploring graph-based connectivity representations, and validating the framework on larger and more diverse datasets. These advancements are expected to further improve model generalizability and clinical applicability.
Acknowledgements
The authors would like to acknowledge the support and resources provided by Karpagam College of Engineering, which were invaluable for the completion of this research.
Author contributions
D.A. conceived the experiment(s). D.A. and S.K. developed the methodology and conducted the experiment(s). S.K. performed the formal analysis and visualized the results. D.A. and S.K. carried out the investigation and curated the data. D.A. and S.K. were responsible for software implementation, validation, and resource management. D.A. prepared the original draft, and S.K. reviewed and edited the manuscript. S.K. supervised the project and handled project administration. D.A. and S.K. acquired funding. All authors reviewed and agreed to the final version of the manuscript.
Data availability
The preprocessed functional data were obtained from the Preprocessed Connectomes Project (PCP) repository (http://preprocessed-connectomes-project.org/abide/), specifically from the ABIDE I release using the Configurable Pipeline for the Analysis of Connectomes (CPAC) with the filt_global preprocessing strategy. The dataset is publicly available for research purposes under the data usage terms of the PCP initiative, and appropriate citations have been followed. The complete code and analysis scripts used to process the data and reproduce the results are publicly available via Zenodo at: https://doi.org/10.5281/zenodo.19563123.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Lord, C. et al. Autism spectrum disorder. Nat. Rev. Dis. Primers. 6, 5. 10.1038/s41572-019-0138-4 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Zwaigenbaum, L. et al. Early intervention for children with autism spectrum disorder under 3 years of age: recommendations for practice and research. Pediatrics136 (Suppl 1), S60–S81. 10.1542/peds.2014-3667E (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Pierce, K. & Courchesne, J. Evidence for a cerebellar role in reduced exploration and stereotyped behavior in autism. Biol. Psychiatry. 49, 655–664. 10.1016/S0006-3223(00)01008-8 (2001). [DOI] [PubMed] [Google Scholar]
- 4.Aarthi, D. & Kannimuthu, S. A comprehensive analysis of autism spectrum disorder using machine learning algorithms: survey. In: International Conference on Power Engineering and Intelligent Systems, Delhi, India. 10.1007/978-981-99-7216-6_20 (2024).
- 5.McKenzie, K. et al. Factors influencing waiting times for diagnosis of autism spectrum disorder in children and adults. Res. Dev. Disabil.45–46, 300–306. 10.1016/j.ridd.2015.07.033 (2015). [DOI] [PubMed] [Google Scholar]
- 6.Schielen, S. J. C., Pilmeyer, J., Aldenkamp, A. P. & Zinger, S. The diagnosis of ASD with MRI: a systematic review and meta-analysis. Transl. Psychiatry. 14, 318. 10.1038/s41398-024-03024-5 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Zhang, S. AI-assisted early screening, diagnosis, and intervention for autism in young children. Front. Psychiatry. 16, 1513809. 10.3389/fpsyt.2025.1513809 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Wang, M., Xu, D., Zhang, L. & Jiang, H. Application of multimodal MRI in the early diagnosis of autism spectrum disorders: a review. Diagnostics13, 3027. 10.3390/diagnostics13193027 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Anagnostou, E. & Taylor, M. J. Review of neuroimaging in autism spectrum disorders: what have we learned and where we go from here. Mol. Autism. 2, 4. 10.1186/2040-2392-2-4 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Alves, C. L. et al. Diagnosis of autism spectrum disorder based on functional brain networks and machine learning. Sci. Rep.13, 8072. 10.1038/s41598-023-34650-6 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Gkintoni, E. et al. Leveraging AI-driven neuroimaging biomarkers for early detection and social function prediction in autism spectrum disorders: a systematic review. Healthcare13, 1776. 10.3390/healthcare13151776 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Santana, C. P. et al. rs-fMRI and machine learning for ASD diagnosis: a systematic review and meta-analysis. Sci. Rep.12, 6030. 10.1038/s41598-022-09821-6 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Varshney, R. K., Katiyar, A. & Johri, P. Hybrid CNN–RNN models for multimodal analysis of autism spectrum disorder neuroimaging. In: 2025 International Conference on Automation and Computation (AUTOCOM), Dehradun, India, pp 155–160. 10.1109/AUTOCOM64127.2025.10956945 (2025).
- 14.Mahmoudi, E. et al. Gamification in mobile apps for children with disabilities: scoping review. JMIR Serious Games. 12, e49029. 10.2196/49029 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Shahini, A. et al. A systematic review for artificial intelligence-driven assistive technologies to support children with neurodevelopmental disorders. Inf. Fusion. 12410.1016/j.inffus.2025.103441 (2025).
- 16.Yang, Z., Zhang, Y., Ning, J., Wang, X. & Wu, Z. Early diagnosis of autism: a review of video-based motion analysis and deep learning techniques. IEEE Access.13, 2903–2928. 10.1109/ACCESS.2024.3523872 (2025). [Google Scholar]
- 17.Trayvick, J. et al. Speech and language patterns in autism: towards natural language processing as a research and clinical tool. Psychiatry Res.340, 116109. 10.1016/j.psychres.2024.116109 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Abraham, A. et al. Deriving reproducible biomarkers from multi-site resting-state data: An Autism-based example. NeuroImage147, 736–745. 10.1016/j.neuroimage.2016.10.045 (2017). [DOI] [PubMed] [Google Scholar]
- 19.Bassett, D. S. & Sporns, O. Network neuroscience. Nat. Neurosci.20, 353–364. 10.1038/nn.4502 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Jung, W., Jeon, E., Kang, E. & Suk, H. I. EAG-RS: A novel explainability-guided ROI-selection framework for ASD diagnosis via inter-regional relation learning. IEEE Trans. Med. Imaging. 43 (4), 1400–1411. 10.1109/TMI.2023.3337362 (2024). [DOI] [PubMed] [Google Scholar]
- 21.Ning, Q. et al. A deep learning method for autism spectrum disorder identification based on interactions of hierarchical brain networks. Behav. Brain Res.452, 114603. 10.1016/j.bbr.2023.114603 (2023). [DOI] [PubMed] [Google Scholar]
- 22.Hazlett, H. C. et al. Early brain development in infants at high risk for autism spectrum disorder. Nature542, 348–351. 10.1038/nature21369 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Bengs, M., Gessert, N. & Schlaefer, A. 4D spatio-temporal deep learning with 4D fMRI data for autism spectrum disorder classification. arXiv preprint. 10.48550/arXiv.2004.10165 (2020). [Google Scholar]
- 24.Sachdeva, J., Mittal, R., Mehta, J., Jain, R. & Ranjan, A. Resolving autism spectrum disorder (ASD) through brain topologies using fMRI dataset with multi-layer perceptron (MLP). Psychiatry Res. Neuroimag.. 343, 111858. 10.1016/j.pscychresns.2024.111858 (2024). [DOI] [PubMed] [Google Scholar]
- 25.Heinsfeld, A. S., Franco, A. R., Craddock, R. C., Buchweitz, A. & Meneguzzi, F. Identification of autism spectrum disorder using deep learning and the ABIDE dataset. NeuroImage Clin.17, 16–23. 10.1016/j.nicl.2017.08.017 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Parisot, S. et al. Disease prediction using graph convolutional networks: Application to Autism Spectrum Disorder and Alzheimer’s disease. Med. Image Anal.48, 117–130. 10.1016/j.media.2018.06.001 (2018). [DOI] [PubMed] [Google Scholar]
- 27.Liu, X., Hasan, M. R., Gedeon, T. & Hossain, M. Z. MADE-for-ASD: A multi-atlas deep ensemble network for diagnosing Autism Spectrum Disorder. Comput. Biol. Med.182, 109083. 10.1016/j.compbiomed.2024.109083 (2024). [DOI] [PubMed] [Google Scholar]
- 28.Yang, S. et al. M3ASD: Integrating multi-atlas and multi-center data via multi-view low-rank graph structure learning for Autism Spectrum Disorder diagnosis. Brain Sci.15 (11), 1136. 10.3390/brainsci15111136 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Jung, W., Jeon, E., Kang, E. & Suk, H. I. EAG-RS: A novel explainability-guided ROI-selection framework for ASD diagnosis via inter-regional relation learning. IEEE Trans. Med. Imaging. 43 (4), 1400–1411. 10.1109/TMI.2023.3337362 (2024). [DOI] [PubMed] [Google Scholar]
- 30.Cilia, F. et al. Computer-aided screening of autism spectrum disorder: eye-tracking study using data visualization and deep learning. JMIR Hum. Factors. 8, e27706. 10.2196/27706 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Kanhirakadavath, M. R. & Chandran, M. S. M. Investigation of eye-tracking scan path as a biomarker for autism screening using machine learning algorithms. Diagnostics12, 518. 10.3390/diagnostics12020518 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Farhat, T. et al. A deep learning-based ensemble for autism spectrum disorder diagnosis using facial images. PLoS One. 20, e0321697. 10.1371/journal.pone.0321697 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
- 33.Anderson, J. T., Stenum, J., Roemmich, R. T. & Wilson, R. B. Validation of markerless video-based gait analysis using pose estimation in toddlers with and without neurodevelopmental disorders. Front. Digit. Health. 7, 1542012. 10.3389/fdgth.2025.1542012 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Jin, L., Cui, H., Zhang, P. & Cai, C. Early diagnostic value of home video-based machine learning in autism spectrum disorder: a meta-analysis. Eur. J. Pediatr.184, 37. 10.1007/s00431-024-05837-4 (2024). [DOI] [PubMed] [Google Scholar]
- 35.Dia, M., Khodabandelou, G., Md Sabri, A. Q. & Othmani, A. Video-based continuous affect recognition of children with autism spectrum disorder using deep learning. Biomed. Signal. Process. Control. 89, 105712. 10.1016/j.bspc.2023.105712 (2024). [Google Scholar]
- 36.Ganjigunte Prakash, V. et al. Video-based real-time assessment and diagnosis of autism spectrum disorder using deep neural networks. Expert Syst.10.1111/exsy.13253 (2023). [Google Scholar]
- 37.Abbas, H. et al. Multi-modular AI approach to streamline autism diagnosis in young children. Sci. Rep.10, 5014. 10.1038/s41598-020-61213-w (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Ma, W., Xu, L., Zhang, H. & Zhang, S. Can natural speech prosody distinguish autism spectrum disorders? a meta-analysis. Behav. Sci.14, 90. 10.3390/bs14020090 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Chi, N. A. et al. Classifying autism from crowdsourced semistructured speech recordings: machine learning model comparison study. JMIR Pediatr. Parent.5, e35406. 10.2196/35406 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.ABIDE I (Autism Brain Imaging Data Exchange). ABIDE I database access. https://fcon_1000.projects.nitrc.org/indi/abide/abide_I.html, (2012).
- 41.Preprocessed Connectomes Project. Preprocessed Connectomes Project – ABIDE Preprocessed Initiative. http://preprocessed-connectomes-project.org/abide/. (2015).
- 42.Cameron Craddock, Y. et al. Yan, Pierre Bellec The Neuro Bureau Preprocessing Initiative: open sharing of preprocessed neuroimaging data and derivatives. Neuroinformatics 2013, Stockholm, Sweden. (2013).
- 43.Sherkatghanad, Z. et al. Automated detection of Autism Spectrum Disorder using a convolutional neural network. Front. Neurosci.13, 1325. 10.3389/fnins.2019.01325 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Eslami, T., Mirjalili, V., Fong, A., Laird, A. R. & Saeed, F. ASD-DiagNet: A hybrid learning approach for detection of Autism Spectrum Disorder using fMRI data. Front. Neuroinform. 13, 70. 10.3389/fninf.2019.00070 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Almuqhim, F. & Saeed, F. ASD-SAENet: A sparse autoencoder and deep-neural network model for detecting Autism Spectrum Disorder (ASD) using fMRI data. Front. Comput. Neurosci.15, 654315. 10.3389/fncom.2021.654315 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The preprocessed functional data were obtained from the Preprocessed Connectomes Project (PCP) repository (http://preprocessed-connectomes-project.org/abide/), specifically from the ABIDE I release using the Configurable Pipeline for the Analysis of Connectomes (CPAC) with the filt_global preprocessing strategy. The dataset is publicly available for research purposes under the data usage terms of the PCP initiative, and appropriate citations have been followed. The complete code and analysis scripts used to process the data and reproduce the results are publicly available via Zenodo at: https://doi.org/10.5281/zenodo.19563123.





























