Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2025 Nov 3;18(11):2210–2230. doi: 10.1002/aur.70135

Artificial Intelligence Networks Combining Histopathology and Machine Learning Can Extract Axon Pathology in Autism Spectrum Disorder

Arash Yazdanbakhsh 1,2,3,, Kim T M Dang 1, Kelvin Kuang 1, Tingru Lian 1, Xuefeng Liu 4, Songlin Xie 4, Basilis Zikopoulos 2,3,4,5,
PMCID: PMC12661275  PMID: 41178535

ABSTRACT

Axon features that underlie the structural and functional organization of cortical pathways have distinct patterns in the brains of neurotypical controls (CTR) compared to individuals with Autism Spectrum Disorder (ASD). However, detailed axon study demands labor‐intensive surveys and time‐consuming analysis of microscopic sections from postmortem human brain tissue, making it challenging to systematically examine large regions of the brain. To address these challenges, we developed an approach that uses machine learning to automatically classify microscopic sections from ASD and CTR brains, while also considering different white matter regions: superficial white matter (SWM), which contains a majority of axons that connect nearby cortical areas, and deep white matter (DWM), which is comprised exclusively of axons that participate in long‐range pathways. The result was a deep neural network that can successfully classify the white matter below the anterior cingulate cortex (ACC) of ASD and CTR groups with 98% accuracy, while also distinguishing between DWM and SWM pathway composition with high average accuracy, up to 80%. Examination of image regions important for network classification and misclassification, through sensitivity maps, along with multidimensional scaling analysis, helped identify key pathological markers of ASD and highlighted the spectrum of ASD heterogeneity and overlaps with neurotypical characteristics. Large datasets that can be used to expand training, validation, and testing of this network have the potential to automate high‐resolution microscopic analysis of postmortem brain tissue, so that it can be used to systematically study white matter across brain regions in health and disease.

Keywords: anterior cingulate cortex, convolutional neural network, deep neural network, long‐range pathways, short‐range pathways, white matter


Summary.

  • Histopathological examination of the brain in autism can reveal key mechanisms of disruption but remains a challenging and time‐consuming process. We developed a reliable, automatic, and significantly faster method to classify histopathological microscopy images of white matter axonal patterns in autism, using AI.

  • The neural network performed well when images were rich‐labeled, based on a combination of pathway range and pathology information, creating better contextual classification and revealing more details about the general patterns of autism.

  • Examination of image regions important for network (mis)classification, helped identify key pathological markers, highlighted the spectrum of autism heterogeneity and overlaps with neurotypical characteristics. Expanding datasets to train and test this network can further facilitate systematic study across brain regions in health and disease.

1. Introduction

Neural communications and connections are altered in Autism Spectrum Disorder (ASD) and several studies have pinpointed changes in axons in a few frontal, temporal, and callosal regions, as a core underlying pathology (e.g., X. B. Liu and Schumann 2014; Wegiel et al. 2018; Zikopoulos and Barbas 2010). In these studies, the most observed axon pathology involved an increase in the relative density of thin axons accompanied by a parallel decrease in the density of thick axons; however, excessive branching due to the expression of growth axon proteins, thinning of the myelin, changes in the inner/outer diameter ratio (g‐ratio), and increased variability of the trajectory of axons have also been reported.

These structural and molecular features of individual axons underlie conduction speed and pathway strength, and ultimately determine the physiology and fidelity of signal transmission, and the integrity of neural communications (Caminiti et al. 2013; Innocenti and Caminiti 2017). In addition, the relative position and size of axons in the white matter can be used as an indicator of their termination in nearby or distant brain areas. The superficial white matter (SWM) mainly includes short‐range connections whereas, the deep white matter (DWM) includes long‐range excitatory pathways, with axons that are significantly thicker than axons found in the SWM, (Herbert et al. 2004; Schmahmann and Pandya 2006). Therefore, a systematic study of the white matter can reveal key structural and functional principles of the organization of brain pathways and mechanisms of disruption in ASD. However, high‐resolution quantitative examination of features of individual axons requires extensive study by experienced anatomists. Such analyzes are labor‐intensive and time‐consuming, rendering this approach not optimal for large‐scale studies that can help identify core ASD network status and likely mechanisms of disruption in communication.

Machine learning may offer an alternative approach for the classification of high‐resolution microscopic images of myelinated axons in white matter pathways that link nearby and distant cortical and subcortical areas. Deep neural networks (DNNs) are tools employed in machine learning that have performed incredibly well with image classification (Krizhevsky et al. 2012). Moreover, DNN and transfer learning are increasingly used to classify medical images, showing the strengths of such an approach for spatial feature detection (Kundu et al. 2024; Shen et al. 2024; Wang et al. 2024; Yadav and Jadhav 2019; reviewed in Chaki and Deshpande 2024; Wen et al. 2024). Among the various DNNs, GoogLeNet is a convolutional neural network that is established as a good prototype with sufficient power for small‐scale datasets (Krizhevsky et al. 2012). Therefore, in this study we developed custom transfer learning processes and protocols to optimize GoogLeNet for distinguishing high‐resolution microscopic images of myelinated axons in white matter pathways that link nearby and distant brain areas in neurotypical controls and individuals with ASD. We focused on the white matter below the anterior cingulate cortex (ACC) that has a key role in attention, social interactions, emotions, and executive control, processes that are affected in autism (Hill 2004). The ACC is consistently affected in ASD, exhibiting hyperactivity during response monitoring and social target detection (Dichter et al. 2009; Thakkar et al. 2008) and desynchronized activity during working memory tasks (Kana et al. 2006), likely due to local over‐connectivity and long‐distance disconnection of ACC pathways (Courchesne and Pierce 2005). To develop and optimize machine learning algorithms we used a large dataset of light and electron microscopy (EM) images of white matter axons from short‐ and long‐range ACC cortical pathways from previous studies that have shown significant differences in control and ASD groups of adults (García‐Cabezas et al. 2018; Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018).

Classification of pathology and white matter depth by our trained network reached high levels of reliability. Quantification and visualization of regions highlighted as crucial for correct and incorrect network classification reaffirmed previous findings and provided new insights, revealing novel features and cues that were not apparent in histopathological analysis and underlie ASD heterogeneity. Our proposed method can be generalized and applied to detect axonal and pathway feature differences across brain areas in the neurotypical brain and in psychiatric disorders with underlying pathology in brain network connectivity.

2. Materials and Methods

2.1. Tissue and Dataset Selection

We used two datasets, consisting of thousands of digital photomicrographs from postmortem human brain sections of the white matter below ACC that were processed and imaged in the Human Systems Neuroscience Laboratory at Boston University (B. Zikopoulos, PI), as described in previous studies (García‐Cabezas et al. 2018; X. Liu et al. 2020; Trutzer et al. 2019; Zikopoulos and Barbas 2010, 2013; Zikopoulos, Garcia‐Cabezas, and Barbas 2018). The first dataset consisted of nanometer‐scale EM images, and the second dataset included micrometer‐scale brightfield optical microscopy images stained with the Nissl stain toluidine blue (Tol‐Blue). The brief description below summarizes key information relevant for tissue selection and imaging. Detailed processing methods can be found in the Supporting Information section. We obtained age‐ and gender‐matched postmortem brain tissue that was immersion fixed in 10% formalin, from 12 individuals, 7 neurotypical controls (CTR) and 5 with ASD (Table 1) from the Autism Tissue Program, the Harvard Brain Tissue Resource Center, the Institute for Basic Research in Developmental Disabilities, the University of Maryland Brain and Tissue Bank, the National Disease Research Interchange (NDRI), Anatomy Gifts Registry, and Autism BrainNet. The diagnosis of autism was based on the Autism Diagnostic Interview‐Revised (ADI‐R). Clinical characteristics, including ADI scores, and other data are summarized in Table 1. Cases were matched for the most part, based on tissue availability; however a larger sample of control cases with a wider age range that included older individuals was used for initial training purposes. The study was approved by the Institutional Review Board of Boston University.

TABLE 1.

Demographics, indices, and scores for each individual case.

Case number Age (years) Sex PMI (h) Cause of death ADI‐R c
S (10) C (8V, 7NV) R (3)
ASD1 30 M 16 Congestive heart failure 26 22 (V) 12
ASD2 30 M 20 Myocardial infarction 22 12 (NV) 12
ASD3 44 M 31 Acute myocardial infarction 26 18 (V), 13 (NV) 6
ASD4 40 F 33 Respiratory arrest b 12 14 (V) 8
ASD5 31 M 99 Shooting 18 14 (V) 6
CTR1 42 M 18 Myocardial infarction NA NA NA
CTR2 36 F 18 Unknown NA NA NA
CTR3 36 M 20 Myocardial infarction NA NA NA
CTR4 67 M 30 Pancreatic cancer NA NA NA
CTR5 58 F 30 Pancreatic cancer NA NA NA
CTR6 a 83 M 24 Unknown NA NA NA
CTR7 30 M 41 Unknown NA NA NA

Note: Generally, shorter PMI correlates with better tissue quality for histopathological studies however, we obtained well‐preserved and appropriately stored tissue with a short PMI (on average 31.6 or 25.5 h, when excluding one case with PMI of 99 h), despite limited tissue availability. Our sample included one ASD case with a PMI of 99 h; however, the mean PMI overall was relatively low and comparable to the mean PMI of cases used in similar studies. Signs of postmortem autolysis were minimal in myelinated axons and the white matter and were restricted to minor dissociation of the myelin sheath or small inclusions in the axolemma in some thick axons, and the typical “fried egg” appearance of oligodendrocytes in postmortem tissue, with a seemingly intact nucleus positioned in a bloated, but empty cytoplasm. ASDx: cases from ASD adults, CTRx: cases from the control group.

Abbreviation: PMI, postmortem interval.

a

Exclusive to EM images, not in training for Tol‐Blue images.

b

Asphyxiation or respiratory problems may result in suboptimal tissue quality.

c

The number next to each criterion of the ADI‐R is the cut‐off score for that criterion. An autism diagnosis is indicated when scores in all three behavioral areas meet or exceed the specified minimum cutoff scores. The three criteria are: social interaction (S), communication and language (C): verbal (V) and nonverbal (NV), restrictive and repetitive behaviors (R).

We used coronal ACC tissue blocks, matched based on the human brain atlas (Mai et al. 2015; von Economo and Koskinas 1925/2008) and additional cytoarchitectonic studies of human prefrontal and cingulate cortex (Palomero‐Gallagher et al. 2018, 2008, 2013; Vogt et al. 2013). Images in each dataset (EM and Tol‐Blue) were labeled based on the population type from which the sections were obtained: neurotypical control group (CTR) versus adults with ASD, and further divided according to the location of origin within the white matter: SWM versus DWM. The resulting four classes were labeled as “Control Deep White Matter” (CTR DWM), “Control Superficial White Matter” (CTR SWM), “Autism Spectrum Disorder Deep White Matter” (ASD DWM), and “Autism Spectrum Disorder Superficial White Matter” (ASD SWM). SWM and DWM images were captured below the gray matter of ACC areas 32, 24 and subgenual cingulate (SGC) area 25. The main body of the anterior part of the cingulum bundle was not included in the DWM analyzed in our studies. The DWM sections were adjacent to the SWM sections at an average distance of 3 mm from the overlying gray matter (~2 mm for SGC regions) (Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018; Zikopoulos, Liu, et al. 2018). However, because axons from the cingulum bundle fan out as they reach their gray matter targets or converge and form bundles as they leave their region of origin in the cingulate, we cannot exclude the presence of some of these axons in our samples.

The EM dataset consisted of over one thousand images (Figure 1) at 2500 × 2500 pixels (10 nm/pixel), that covered and sampled an average white matter surface area of 0.3 mm2/case in each of the groups (ASD DWM, ASD SWM, CTR DWM, CTR SWM). The optical microscopy dataset consisted of hundreds of images (Figure 2) at 4080 × 3072 pixels (42.78 nm/pixel), that covered and sampled an average white matter surface area of 1.6 mm2/case in each of the groups (ASD DWM, ASD SWM, CTR DWM, CTR SWM).

FIGURE 1.

FIGURE 1

Representative EM images of ACC and SGC white matter. (A–D) Images from the DWM of 2 CTR (A, B) and 2 ASD (C, D) cases. (E–H) Images from the SWM of 2 CTR (E, F) and 2 ASD (G, H) cases. Scale bar in H applies to all panels.

FIGURE 2.

FIGURE 2

Representative light microscopy images from Toluidine Blue‐stained sections of ACC and SGC white matter. (A–D) Images from the DWM of 2 CTR (A, B) and 2 ASD (C, D) cases. (E–H) Images from the SWM of 2 CTR (E, F) and 2 ASD (G, H) cases. Scale bar in H applies to all panels.

2.2. Image Preprocessing Methods

As network training typically benefits from a large dataset, we considered two image preprocessing methods that would increase the size of the dataset: the N × N window‐tiling method, and the sliding window method. Both methods were applied to determine the optimal sub‐image size for the best network accuracy and implemented by nested “for” loops in MATLAB 2023a. The N × N window‐tiling method involved cutting the original image into smaller sub‐images by splitting each dimension by N, creating N 2 images. For example, if N = 2, one original image would generate four sub‐images. By considering different N values, each classification accuracy was evaluated and compared with others, and the best‐performing N value with its corresponding sub‐image size was determined. After implementation of the window‐tiling method, for EM we had an average of 2500 images (ASD) and 2000 images (CTR), while there were 7380 images per class for the Tol‐Blue slices.

The second image preprocessing method we considered was the sliding window method. Similarly to the sub‐imaging method, the larger image would be divided into smaller sub‐images, but a sliding window was applied, resulting in overlaps within the sub‐images (Figure 3). For the EM dataset, the sliding window moved 500 pixels (1/4 the sub‐image size) for each sub‐image in both the x‐ and y‐axes. For the optical microscopy dataset, the sliding window moved 89 pixels (1/5 the sub‐image size) for each sub‐image in both the x‐ and y‐axes.

FIGURE 3.

FIGURE 3

Specifications and preprocessing of microscopic section images. (A) An example of a source EM section image, that can be digitally cropped to generate subsections to be used as DNN input; note the color correspondence of square outlines of cropped areas before and after the arrow to visualize subsections from the source section. (B) An illustration of the sliding window method used for Tol‐Blue images; note the stride of windows (sequence of colored square outlines) and the resultant sections to be input to DNN. Again, the square outline color correspondence before and after the arrow shows corresponding subsections.

EM images were first cropped from 2500 × 2500 to 2000 × 2000 pixels and then downsampled/compressed to 224 × 224 pixels whereas, the Tol‐Blue images were cropped from 4080 × 3072 to 448 × 448 pixels and then compressed to 224 × 224. Images that were used as inputs for CNN transfer learning, training, validation, and testing were 224 × 224 pixels, as this was the pixel size of the images used for the original training of the Inception V1 (GoogLeNet) (Christian Szegedy et al. 2014) and ResNet‐18/50 (He et al. 2015) networks. In addition, for the Tol‐Blue optical microscopy dataset, the images were first converted to grayscale to remove any background features that the color of the dye may have introduced. Due to the varying contrasts within the slices, contrast‐limited adaptive histogram equalization was also applied to create a more uniform dataset by utilizing Image Processing and Computer Vision, and Statistics & Machine Learning toolboxes of MATLAB 2023a. The images were also normalized to the values between [0, 1] instead of [0, 255] for training.

To avoid overfitting in the trained network, we employed a data‐splitting method that would ensure a complete separation across cases between the training and testing datasets. The EM images and the optical microscopy images were separated differently due to their variances in case and image availability. Specifically, for the EM images we used two different methods of data distribution between training, validation, and testing datasets (Tables 2 and S1). The first method separated the training/validation data from the testing dataset, with one case (~20% of the dataset) going into the testing dataset, while the rest of the cases were split 7:3 to the training/validation set. The second method separated the validation/testing dataset from the training dataset, with one case being split 50:50 to the validation/testing dataset, while the rest of the cases went to the training set.

TABLE 2.

Rank‐based training on different case combinations.

Rank a Training data Validation data Test data 4 class accuracy Pathology accuracy Depth accuracy
A. EM
1 ASD2, ASD3, ASD4, ASD5, CTR1, CTR2, CTR3, CTR4, CTR5 ASD1, CTR6 b ASD1, CTR6 b 52.65% 90.88% 57.21%
7 ASD2, ASD3, ASD4, ASD5, CTR1, CTR2, CTR3, CTR5, CTR6 b ASD2, ASD3, ASD4, ASD5, CTR1, CTR2, CTR3, CTR5, CTR6 b ASD1, CTR4 58.57% 72.30% 76.98%
9 ASD2, ASD3, ASD4, ASD5, CTR1, CTR2, CTR3, CTR5, CTR6 b ASD1, CTR4 ASD1, CTR4 52.85% 69.78% 79.75%
B. Tol‐Blue—3 case training
1 ASD2, ASD3, ASD4, CTR2, CTR3, CTR4 ASD5, CTR5 ASD5, CTR5 67.52% 97.68% 69.13%
C. Tol‐Blue—2 case training
1 ASD4, ASD5, CTR3, CTR4 ASD3, CTR5 ASD2, CTR2 48.94% 60.97% 48.94%
2 ASD3, ASD4, CTR3, CTR5 ASD5, CTR4 ASD2, CTR2 49.56% 59.44% 49.56%

Note: Ranking is based on pathology classification accuracy (descending). This table includes only the network models with the highest accuracy in one of each of the three testing categories: 4 Class Accuracy; Pathology Accuracy; Depth Accuracy. The full ranking is shown in Table S1. (A) Results for EM images. (B) Results for 3 case training (3 CTR and 3 ASD) on Tol‐Blue images. (C) Results for 2 case training (2 CTR and 2 ASD) on Tol‐Blue images. Our exhaustive training and testing with different combinations of nonoverlapping cases, showed that with small datasets, in this case limited number of cases and images, the best approach was to train separate models for optimal classification of pathology only, or white matter depth only, or combinations of these (best classification accuracies in bold font). Note that there was no overlap of cases and images between training and testing (also shown in full detail in Table S1). In addition, there was no overlap of images between training and validation, even when training and validation sets came from the same cases (e.g., EM rank‐based training session 7).

a

Ranking of networks is based on best pathology classification accuracy.

b

CTR6 case was used exclusively for EM analysis.

2.3. Model Training, Validation, and Testing

To increase the generalizability and reliability of the model we ensured that each training and testing set included images from different cases and that there were no shared sub‐images in either dataset. Therefore, we ensured that there was no overlap of cases and images between training and testing. In addition, there was no overlap of images between training and validation, even when the training and validation sets came from the same cases. We addressed the variability in the number of images across cases by employing both undersampling and oversampling techniques to balance the training dataset for each class. We utilized undersampling by identifying the class with the fewest images and limiting the training set for each class to that minimum number through random selection. Conversely, for oversampling, rather than reducing the number of images for larger classes, we opted to duplicate images in smaller classes, ensuring they reached a comparable image count to the larger classes. Following experimentation with both methods, the results indicated that undersampling produced superior outcomes. This is attributed to its ability to mitigate the potential overfitting effects that may arise from oversampling the same set of images.

2.3.1. Transfer Learning

Due to the small size of the dataset, we performed transfer learning, which is the process of using a pretrained model, trained on a large and general dataset, to attempt to generalize to new data, in this case, brain microscopy images at the cellular and subcellular level. We tested different pretrained models including AlexNet (Krizhevsky et al. 2012), SqueezeNet (Iandola et al. 2016), EfficientNet‐b0 (M. Tan and Le 2019), Inception V1 (GoogLeNet) (Christian Szegedy et al. 2014), ResNet‐18/50 (He et al. 2015), and VGG‐16/19 (Simonyan and Zisserman 2014). Inception v1 and ResNet‐18 resulted in the best networks. The optimal hyperparameters were found using stochastic gradient descent with momentum, a learning rate of 0.001 with a drop factor of 0.1 every 2 epochs, 6 epochs, a mini‐batch size of 128, and L2 regularization lambda of 0.0001. We employed early stopping if the loss did not improve after an epoch.

Pretrained models (such as Inception v1 and ResNet‐18) can be fine‐tuned or used for feature extraction. Instead of re‐training the entire model, fine‐tuning involves freezing the initial layers of a model while leaving the top layers, including a newly added classifier layer, unfrozen. When a layer is frozen, the weights are not updated when training on the new data. Earlier layers of a model focus on lower‐level features (such as lines and edges), while the top layers extract higher‐order features. The top layers are left unfrozen since higher‐order features have more relevance when training on new data, while the lower‐level features are more common. Feature extraction is the process of using a pretrained network to extract the features of new data and training a new classifier from those features. Since most pretrained CNNs perform well when using raw images as the input, training a new classifier on the extracted feature vectors is skipped and training is performed directly by fine‐tuning the top layers.

2.3.2. Model Optimization

When training a model, there are multiple optimizers to choose from. Optimizers are algorithms that are used to update the weights of a neural network to minimize the loss (the penalty for incorrect classifications). In this study, we tested two of the most commonly used optimizers: SGDM (stochastic gradient descent with momentum) and Adam (adaptive movement estimation). SGDM, defined by the equation:

θt+1=θtαLθt+γθtθt1

in which: θ= parameters (weights), t= iteration number, α= learning rate, L= loss function, γ= momentum, updates each weight to minimize the loss function by computing the gradient of the loss function and updates the weight using a random mini‐batch. Since SGD (stochastic gradient descent) can oscillate toward the optimum, SGDM implements a momentum term to reduce the oscillation. Unlike SGDM, Adam differs by using learning rates that differ by parameter and can automatically adapt to the loss function being optimized. Adam maintains an average of both the parameter gradients (1) and the square of the parameter gradients (2):

mt=β1mt1+1β1Lθt (1)
vt=β2vt1+1β2Lθt2 (2)

in which: mt= average of parameter gradients, vt= average of square parameter gradients, t= iteration, β1= gradient decay, β2= square gradient decay, L= loss function, θ = parameter. Adam uses the moving averages to update the network parameters:

θt+1=θtαmtvt+ϵ

where θ= parameter (weights), t= iteration, α= learning rate, mt= average of parameter gradients, vt= average of square parameter gradients, ϵ= division by zero constant. Compared to SGDM, Adam can update each parameter based on the gradient's history. Adam is a faster algorithm compared to SGDM, but it is more susceptible to noise, which may lead to suboptimal performance.

A model must be given the optimal hyperparameters, so that training and testing accuracy converge smoothly, resulting in fitting and avoiding overfitting; therefore resulting in the model learning the key general patterns instead of memorizing (overfitting). Hyperparameters (learning rate, momentum, epochs, mini‐batch, etc.) control how the model trains and determine the parameters of the model. In order to optimize the hyperparameters, we used two algorithms: random search and Bayesian optimization. Random search is the process of generating and evaluating random parameters, in this case to minimize loss. Random search is an uninformed algorithm, as it has no knowledge of the previous results. Unlike random search, Bayesian optimization is an informed algorithm. By using the previous results, Bayesian optimization uses a probabilistic model to estimate the next parameters. We used both of these methods to find initial hyperparameters and performed manual tuning to achieve the final hyperparameters.

2.3.3. Overfitting Prevention

Overfitting is a problem in machine learning where a model learns the background noise in the training set and is unable to generalize to the new data in the testing set. The methods we used to combat overfitting are: L2 regularization, dropout, reducing model complexity, and freezing layers. L2 regularization penalizes the model if the weights become too large and encourages the model to have small and a more even distribution in its weights. L2 regularization is included in the hyperparameters that can be tuned to minimize the loss (difference between the predicted value by the model and the original value given in the dataset); thereby the networks generalize better to new data (Pedregosa et al. 2011). Dropout is another regularization method that can be implemented to different layers within a CNN. By implementing dropout, a different node within a CNN will randomly be disabled (or “dropped”) with probability p. Dropout helps combat overfitting by forcing the model to learn more robust features, which can improve generalization. If a model picks up the background noise and features of the data, it can be a sign that the model is too complex for the task at hand, and removing layers to create a simpler model can help with overfitting. Freezing the layers can also help reduce overfitting by limiting the number of parameters that are updated when training, also reducing the model complexity.

2.4. Analysis, Statistics, and Visualization of Results

2.4.1. Evaluation of Model Performance: k‐Fold Cross Validation

We used several independent methods to estimate the accuracy of the model and analyze findings. Cross‐validation is a general method used to determine the accuracy of a trained machine learning model on unseen data. In this study, we used k‐fold cross‐validation, which provides a more reliable performance evaluation compared to a train‐test split (Burman 1989). For k‐fold cross‐validation, we divided the original dataset into k equal‐sized partitions, or “folds.” The model was trained k times, each time using k‐1 folds as the training data, and the remaining fold as the testing data. This ensured that each fold was used once for testing. In our design, each fold was a case, and different cases were rotated between the training and testing sets. The performance of each model was evaluated using metrics such as accuracy or other precision and recall metrics. By averaging the results from each model, an estimation of the overall performance of the model was obtained. By using k‐fold cross‐validation, the model was trained and evaluated on different subsets of the data, facilitating assessment of the generalizability of the model.

2.4.2. Evaluation of Model Performance: Classification Confusion Matrices

We additionally used confusion matrices to evaluate the accuracy of the network after training. In confusion matrices, the rows display the true classes, while the columns display the predicted classes, or what the network predicts which class the images belong to. For multiclass confusion matrices, the on‐diagonal displays the correctly classified images, also known as true positives. The values in a column for a specific class that is not the true positive reflect the false positives for that class, the values in the row that is not the true positive reflect the false negatives for that class, and all other values not in the corresponding row and column for that class reflect the true negatives. We additionally combined the four original classes (ASD SWM, ASD DWM, CTR SWM, and CTR DWM) into ASD versus CTR and DWM versus SWM to analyze the network performance for classification of pathology and white matter region/depth, respectively.

2.4.3. Evaluation of Model Performance: Estimation of Precision and Recall Metrics

To take into account false positives and negatives in our accuracy measures we estimated two additional evaluation metrics, precision and recall, which in image classification tasks provide in‐depth analysis of the performance of a machine learning model (Powers 2008). For multiclass classification, precision, also known as positive predictive value, is defined as the fraction of correct predictions (true positives) over all predictions (true positives and false positives) for a single class. A higher precision signifies a low prevalence of false positives, and when the network predicts a certain class, it is likely to be correctly classified. Recall, also known as sensitivity or true positive rate (TPR), is the fraction of true positives over all actual positives (true positives and false negatives). A higher recall score signifies a low prevalence of false negatives (misses), and that the network can accurately identify instances of true positives. By calculating the harmonic mean of both precision and recall, the F1 score is obtained. Compared to standard accuracy, the F1 score provides a more balanced metric, taking into account false positives and negatives.

2.4.4. Evaluation of Model Performance: Receiver Operating Characteristic Curves

A second performance evaluation metric that we implemented is the receiver operating characteristic curve (ROC curve), which shows the performance of a neural network at different classification thresholds/criteria. An ROC curve plots the TPR versus false positive rate (FPR), and by altering the classification threshold, the TPR and FPR can increase or decrease, depending on the classification threshold. The best operating point can be chosen based on the desired TPR and FPR. Another aspect of an ROC curve is the area under the curve (AUC). The AUC provides an aggregate measure of performance across all possible classification thresholds and represents the network's ability to differentiate between true and false positives. The closer the AUC is to 1.0, the better the model's ability to separate classes from each other.

2.4.5. Image Sensitivity Maps

To further analyze and visualize the results we generated image sensitivity maps using the following two approaches: Occlusion sensitivity maps, also known as saliency maps, provide insights into how the network creates its predictions by identifying the salient regions of the image that positively and negatively contribute to the network's classification. Occlusion sensitivity maps operate by sliding an occluding mask across the input data, resulting in a change in the classification score at each mask location. This process enables the identification of specific areas within the input data that exert the most influence on the classification score. In addition to experimenting with the default size and stride for the occluding mask (with a mask size of 20% of the input size and a stride of 10% of the input size), we explored various combinations of stride and size to determine the optimal patch for highlighting axonal details.

Gradient class activation mapping (Grad‐CAM) is a second approach utilized to provide insight into the regions of interest for network predictions. Grad‐CAM, which is based on the CAM technique, determines the importance of each neuron in a network prediction by considering the gradients of each target classification flowing through the deep network. This results in a localization map highlighting the regions of interest for the desired class. Compared to occlusion sensitivity, Grad‐CAM provides a more general overview of the features that contribute to the class, while occlusion sensitivity is more specific to what specific areas the network identified to make its prediction. Grad‐CAM can be used in similar ways as occlusion sensitivity, but these two image sensitivity maps are best used complementarily to find generalizations and specializations, respectively. Using both approaches, we analyzed the overlaps and lack of overlaps in the regions of interest of the model versus significant features of ACC white matter from prior neuroanatomical studies (García‐Cabezas et al. 2018; Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018). Our aim was to use image sensitivity maps to decipher the aspects of network focus, so that we can then (a) identify novel features and metrics that have not been studied extensively before or have not reached significance in prior comparisons and, (b) identify features that were not emphasized by the network but featured prominently in prior neuroanatomical studies.

When graphing image sensitivity images, we also included the probability predicted by the network for that class to show the level of certainty the network had for the classification of a certain image. In addition to producing heat maps for the specific class (e.g., ASD DWM, ASD SWM, etc.), we also included the combined grad‐cam for each pathology, such as combining ASD DWM and ASD SWM for a complete ASD representation. Compared to other saliency map techniques such as occlusion sensitivity, Grad‐CAM provides a more global and broader outlining view of regions that contribute most toward classification, while other position‐based saliency maps such as occlusion sensitivity provide a more local view (Figure S1).

2.4.6. Generation of DeepDream Images

To better visualize the data and highlight patterns that were characteristic for each group and played a key role for correct classifications we generated DeepDream images. DeepDream, a computer vision algorithm created by Google in 2015, enhances the features found in images and creates visualizations of the features extracted by a neural network for a desired classification (Mordvintsev et al. 2015). After a network is trained, the DeepDream algorithm alters an image consisting of random noise to maximize “loss,” or the sum of activations for a chosen layer of a neural network to excite the layer for a desired class. Earlier layers may produce DeepDream images that consist of simpler patterns such as edges or colors, while later layers are more detailed and include sophisticated features. When training a network, loss is typically minimized through gradient descent, but the DeepDream algorithm maximizes loss through gradient ascent for a specified class. By generating DeepDream images from the training data, the enhanced features can be noted for each class, cross‐validated with image sensitivity maps, and can be used as guidelines for future classifications.

2.4.7. Principal Component Analysis

To extract and analyze relevant image features and metrics for each classification of the CNN we applied dimension reduction techniques, typically used in conjunction with machine learning approaches (Ranaut et al. 2024; Wen et al. 2024). For each convolution layer in a neural network, an activation function is used to convert the input values into an output. In terms of image classification, the output is typically a value that corresponds to a specific class (Krizhevsky et al. 2012; C. Szegedy et al. 2015). The fully connected layer, which receives inputs from all its previous layers, performs a weighted summation on the class‐specific activations, and feeds the values into a SoftMax activation function that normalizes the values into probabilities to determine the final output class. By utilizing the layer activations from the fully connected layer, each input, represented by a layer activation vector, can be analyzed for its similarities and uniqueness for each class. For multiclass classification tasks, the activation vectors can contain N dimensions, but the relevant information is contained in a smaller M dimension. To extract the relevant M dimensions, dimension reduction techniques are applied, and in this study, we specifically used principal component analysis (PCA), to better analyze and visualize the data. PCA is an optimal dimension reduction technique that captures the maximum variance in the data through singular value decomposition (SVD). Through this process, features that provide no information or negligible information are less represented, and the best low‐dimensional approximation of the data is constructed. In order to perform PCA, the original data matrix was normalized, the covariance matrix was computed and then decomposed into its eigenvalues. Given an m × n matrix M , the standardized matrix Z was computed as:

Z=Mμσ

where 𝜇 is the mean and σ is the standard deviation. The covariance matrix C of Z is given by:

C=1n1ZTZ

where n is the number of observations in the data. Eigenvalue decomposition of C yields:

C=VDVT

where V is the matrix of eigenvectors and D is the diagonal matrix of eigenvalues. The eigenvectors were ordered by their corresponding eigenvalue in descending order, which can be used to determine the importance of each principal component. After performing eigenvalue decomposition, the loadings were also calculated, which indicate how much each variable contributes to each principal component and helps understand the meaning of the principal components in terms of the original variables. After eigenvalue decomposition, the loadings matrix L was calculated by:

L=VD1/2

where D 1/2 is the diagonal matrix, whose diagonal elements are the square roots of the eigenvalues. Each column of the loadings matrix L corresponds to a principal component, and each row corresponds to one of the original variables. The magnitude of a loading indicates the importance of the corresponding variable to the principal component, and the sign of a loading indicates the direction of the relationship between the variable and the principal component. Compared to other dimension reduction techniques such as MDS, PCA does not use the Euclidean distance of activations, but the value of the activations from the layers themselves. PCA increases the interpretability of higher dimensional data while preserving the key features of the data as specified by the singular values obtained from SVD.

2.5. Code/Software

All codes and analyzes are deposited and available in https://github.com/kkuang0/ASD‐ACC‐Transfer‐Net (public access).

3. Results

3.1. DNNs Perform Well With Pathology Classification and Single Out ASD SWM

We partitioned the dataset into three subsets: a training set for network training, a validation set for periodic progress assessment during training, and a testing set to evaluate the final network's accuracy. This data‐splitting procedure was repeated multiple times, each time using different case combinations for training, validation, or testing. Tables 2 and S1 present the training results for all combinations, ranked from highest to lowest pathology classification accuracy.

The highest‐performing network for pathology distinction using EM images achieved an accuracy of 52.65% across four classes (Figure 4A.i). Notably, focusing solely on pathology classification resulted in significantly improved performance, with the highest accuracy reaching 90.88% for EM images (Figure 4A.ii). For Tol‐Blue optical microscopy images, the best‐performing network for pathology distinction achieved an accuracy of 67.52% across four classes (Figure 4B.i) and 97.68% for pathology (Figure 4B.ii). Conversely, classification accuracy for DWM versus SWM was lower, with 57.21% for EM images (Figure 4A.iii) and 69.13% for Tol‐Blue images (Figure 4B.iii). Therefore, the major source of the low 4 class accuracy stems from the DWM versus SWM confusion.

FIGURE 4.

FIGURE 4

Confusion matrices and ROC curves for model classification. (A) Classification of EM images. (B) Classification of Tol‐Blue microscopy images. For both panels (A, B): (i) 4 class confusion matrix of pathology and white matter region, (ii) white matter depth independent pathology matrix, (iii) pathology independent white matter depth matrix, (iv) 4‐class ROC, (v) white matter depth independent ROC, (vi) pathology independent white matter depth ROC. Confusion matrices show the true classes compared to the predicted classes. For 4‐class confusion matrices (A.i, B.i), the column summaries correspond to precision, and the row summaries correspond to recall for the corresponding classes. For 2‐class confusion matrices (A.ii, A.iii, B.ii, B.iii), the confusion matrix cells (from left to right, top to bottom) correspond to true positive (TP), false negative (FN), and false positive (FP), and true negative (TN) (with respect to ASD for panels marked [ii], and to DWM for panels marked [iii]). An ROC curve plots the true positive rate (TPR) against the false positive rate (FPR) and shows the performance of a classification model at all classification thresholds. The model operating points represent the threshold that minimizes the FPR and maximizes the TPR.

For EM images, the classes that had the highest accuracies were ASD DWM and CTR DWM, 88.8% and 96.5% respectively (Figure 4A.i) Conversely, the class with the most misclassifications was ASD SWM with most of them being misclassified as ASD DWM, with an accuracy of only 14.7%. This aligns with the characteristics of patients with ASD, whose SWM is typically the most variable (Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018; Zikopoulos, Liu, et al. 2018).

For Tol‐Blue images, the network perfectly classified ASD SWM, but also classified 96.57% of ASD DWM slices as ASD SWM (Figure 4B.i). Of all slices classified as ASD, 98.29% were classified as ASD SWM. This suggests that the network mainly learned the features of ASD SWM and generalized those features as ASD. The network could distinguish between different CTR white matter depths with similar accuracies: 83.9% for DWM and 82.8% for SWM.

3.2. Dimension Reduction Graphs Further Support the Classification Results

To visualize the dataset in three‐dimensional space and observe the variance among the data points, we extracted the feature vectors from the last fully connected layer (see Methodology) to generate 3D PCA graphs from the best‐performing network for EM and Tol‐Blue images (Figure 5).

FIGURE 5.

FIGURE 5

PCA biplots of EM and Tol‐Blue microscopy images. (A) EM analysis. (B) Tol‐Blue analysis. (i) 4‐class, (ii) ASD versus CTR, (iii) DWM versus SWM, (iv) ASD DWM versus ASD SWM, (v) ASD DWM versus CTR DWM, (vi) ASD SWM versus CTR SWM, (vii) CTR DWM versus CTR SWM. The axes correspond respectively to the first, second, and third principal component. PCA loadings describe how much each variable contributes to a particular principal component and the variance explained by that principal component. Plots for pooled categories, in (ii) and (iii), highlight that the higher accuracy of the classification for pathology, as opposed to white matter depth, showing that the ASD versus CTR data points were significantly more dispersed and distinct, whereas the white matter depth data points were more interspersed and less distinguishable.

For EM images, consistent with the confusion matrices, the PCA plots clearly distinguished between the ASD data points and the CTR data points (Figure 5A.i, ii), when viewing the first three principal components. On the other hand, the data points for DWM versus SWM were more mixed in with one another, indicating a higher confusion for white matter layer classification (Figure 5A.iii) The loading vectors for the classes also showed a similar trend as ASD and CTR vectors appeared further apart and more distinct (angle > 90°, Figure 5A.ii) compared to the SWM and DWM vectors (angle < 90°, Figure 5A.iii), suggesting differences in the relationship between each variable and the principal component. That is because the magnitude of a loading indicates the importance of the corresponding variable to the principal component, and the sign of a loading indicates the direction of the relationship between the variable and the principal component.

For Tol‐Blue images, the PCA plots showed similar results as the EM images, with a clear separation between ASD and CTR, but less distinction between the white matter depths (Figure 5B.i–iii). When comparing the loadings for pathology (Figure 5B.ii) and white matter depth (Figure 5B.iii), the ASD and CTR loadings had an angle of 99.823° (1.742 rad), while the DWM and SWM loadings had an angle of 94.722° (1.653 rad), which is consistent with better white matter depth separation in Tol‐Blue compared to EM images.

Pairwise plotting from 2 of 4 total classes also reveals similar results, with the most confusion coming from the DWM versus SWM pairs (ASD DWM vs. ASD SWM and CTR DWM vs. CTR SWM) (Figure 5A.iv, v, 5B.iv, v). Figure 5A,B.vi, vii clearly shows the distinction between the ASD and CTR groups, whether it is ASD DWM versus CTR DWM, or ASD SWM versus CTR SWM as the data points in each respective group are spread apart from left to right.

3.3. PCA Emphasizes Variance and Correlations

To understand the importance of each principal component, we also estimated the percentage of total variance explained by each component. In EM images, the first principal component captured 70.01% of the explained variance, while the second principal component added in another 14.78% (Figure 6A.ii). Similarly, in the Tol‐Blue slices, we observed that 70.1% of the explained variance was captured by the first principal component and an additional 25.3% by the second principal component (Figure 6B.ii), which indicates that the variability of the data was successfully captured by the first two principal components. We chose the third principal component to be the cutoff to maximally contain as much explained variance (94.11% and 98.8% for EM and Tol‐Blue, respectively), while still reducing the dimensions to facilitate interpretation.

FIGURE 6.

FIGURE 6

Scree and loading plots after PCA of EM (A) and Tol‐Blue (B) microscopy images. (i) Eigenvalues of each principal component, (ii) explained variance by each principal component and cumulative explained variance by all principal components, (iii) coefficients of each class contributing to each principal component.

The coefficients for each principal component, known as loadings, also revealed crucial details of how the network viewed the data. Each loading vector represents how much each variable contributes to a principal component, while the direction and magnitude of each vector indicate correlations and magnitudes between each class and principal component. The different coefficients of each loading indicate how each class contributes to the component. In the first principal component loading, ASD classes and CTR classes had different signs (Figures 6A.iii and 7B.iii). For EM images, both ASD DWM and ASD SWM had positive coefficients: 0.7279 and 0.1523, while their corresponding CTR classes were negative: −0.4256 and −0.5155 (Figure 6A.iii). We saw the inverse for Tol‐Blue slices, where both ASD classes had negative coefficients, and the CTR classes had positive coefficients (Figure 6B.iii). This indicates that the first loading feature corresponded to the separation between ASD and CTR. This also explains why there is such a clear distinction between the ASD and CTR classes in the scatter plots (Figure 5), as the variance between ASD and CTR is mainly captured. The distinction between DWM and SWM regions was also reflected in the loadings. In the second principal component loadings for EM images (Figure 6A.iii), the DWM regions were positive (0.2951 and 0.8865), while the SWM regions had negative values (−0.4257 and −0.5155). The same was true for Tol‐Blue images for its third principal component (Figure 6B.iii), as Tol‐Blue's loading had negative coefficients for the DWM regions, and positive coefficients for the SWM regions, indicating that this feature was responsible for the distinction between white matter regions. In the third principal component loading for EM, all the coefficients were positive, with a tendency for the loadings of DWM regions to be closer to zero (0.2490 and 0.1810), while the loadings from SWM were larger (0.8384 and 0.4498); the sizable difference between ASD SWM's loading compared to other classes points out the confusion about ASD SWM. This is consistent with the confusion matrix as the network clearly struggled to identify ASD SWM correctly and mostly misclassified it as ASD DWM (Figure 4A.i). There was a reverse confusion as well in Tol‐Blue images, where the network consistently confused ASD DWM as ASD SWM (Figure 4B.i). This can be explained by the second principal component, as the coefficient for ASD SWM is significantly smaller compared to the other classes, indicating that the second principal component mainly differentiated the other classes from ASD SWM (Figure 6B.iii). We also found that the explained variance for the principal components responsible for DWM and SWM separation only covered 14.78% for EM, and 3.3% for Tol‐Blue (Figures 6A.ii and 7B.ii). This can explain why the separation between DWM and SWM was less clear compared to pathology classification.

FIGURE 7.

FIGURE 7

Image sensitivity maps of correctly classified EM images. (A–D) Representative examples of EM images and overlaid heat maps from each of the four classes: CTR DWM (A), CTR SWM (B), ASD DWM (C), ASD SWM (D). For each class we used Grad‐CAM to generate eight heatmaps, representing hotspots and their relative weights (probability) for correct and incorrect class assignment: (i) ASD DWM, (ii) ASD SWM, (iii) CTR DWM, (iv) CTR SWM, (v) ASD, (vi) CTR, (vii) DWM, (viii) SWM. The title of each heatmap shows the sum of model neuron activity from the fully connected layer, as well as the SoftMax probability. Heatmaps of the correct corresponding class highlight areas of the image that contributed most toward the overall classification. Heatmaps of the other classes provide a comparative visualization of how different features were weighted. This aids in identifying specific patterns and features that distinguish each individual class, enhancing the interpretability and transparency of the network's performance.

3.4. Heat Maps Reveal Crucial Areas for Classification

We used heat maps produced by occlusion sensitivity and Grad‐CAM approaches to analyze and highlight the overlaps or lack thereof in the regions of interest that weighed most for the model classification output and compared with significant features and pathology of ACC white matter from prior neuroanatomical studies (García‐Cabezas et al. 2018; Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018). Grad‐CAM overall provided a more global and broader outlining view of regions that contributed most toward classification that was easier to interpret, whereas occlusion sensitivity provided a more local view that was in many cases too narrow and difficult to interpret (compare Figures 7 and S1). Grad‐CAM‐generated saliency maps for EM (Figure 7) and Tol‐Blue images (Figure 8) consistently highlighted the saliency of axon density, size, and trajectory variability as key features for correct classification. Importantly, hotspots that included cells (mainly nuclei of glia) and relatively uniform axon regions were typically salient in incorrectly classified images (Figure S2).

FIGURE 8.

FIGURE 8

Image sensitivity maps of correctly classified Tol‐Blue microscopy images. (A–D) Representative examples of Tol‐Blue images and overlaid heat maps from each of the four classes: CTR DWM (A), CTR SWM (B), ASD DWM (C), ASD SWM (D). For each class we used Grad‐CAM to generate eight heatmaps, representing hotspots and their relative weights (probability) for correct and incorrect class assignment: (i) ASD DWM, (ii) ASD SWM, (iii) CTR DWM, (iv) CTR SWM, (v) ASD, (vi) CTR, (vii) DWM, (viii) SWM. The title of each heatmap shows the sum of model neuron activity from the fully connected layer, as well as the SoftMax probability. Heatmaps of the correct corresponding class highlight areas of the image that contributed most toward the overall classification. Heatmaps of the other classes provide a comparative visualization of how different features were weighted. This aids in identifying specific patterns and features that distinguish each individual class, enhancing the interpretability and transparency of the network's performance.

3.5. DeepDream Images Generalize the Extracted Features From Each Class

DeepDream is an iterative process that alters an image of random noise to maximally increase the confidence level for the specified output class (Mordvintsev et al. 2015). For both EM and Tol‐Blue images, the DeepDream images for DWM regions of both ASD and CTR (Figure 9A.i, iii and 9B.i, iii) showed distinct axon‐like bundles and clustering, which is a key defining feature of DWM regions. On the other hand, the SWM DeepDream images appeared more convoluted, indicating that SWM regions contained a variety of axonal profiles with heterogeneous trajectories, including cross‐oriented and perpendicular axons (Figure 9A.ii, iv and 9B.ii, iv). When comparing ASD and CTR images, the CTR images showed higher axonal orientational consistency in both DWM and SWM regions compared to the respective ASD regions. Specifically, the CTR DWM image showed more organized distribution patterns of axons, while the ASD DWM image contained higher randomness of axon distribution. When comparing the ASD DeepDream images for EM and Tol‐Blue, we noticed that the EM DeepDream images were characterized by a checkered (square‐like) pattern, while the Tol‐Blue images were characterized by wavy patterns. This could reflect patterns observed in the EM and Tol‐Blue Grad‐CAMs for ASD SWM (see Figures 7 and 8), where the axons in the areas highlighted by the heatmaps in EM formed a more rectangular structure, while the Tol‐Blue Grad‐CAMs mostly highlighted parallel axons, creating the distinct bundles.

FIGURE 9.

FIGURE 9

DeepDream images can generalize the salient extracted features from each class (A) EM and (B) Tol‐Blue DeepDream images constructed from the networks for each class: (i) ASD DWM, (ii) ASD SWM, (iii) CTR DWM, (iv) CTR SWM. EM DeepDream images showed the network's maximally amplified activations for EM features, highlighting the textures the network extracted for each class. Tol‐Blue DeepDream images similarly showed the network's maximally amplified activations for Tol‐Blue features, highlighting more of the structures as well as textures the network extracted for each class. In the DWM sections for both EM and Tol‐Blue, the DeepDream images created distinct axon bundles, while the SWM images contained sparse, scattered axons of different orientations. These DeepDream images offer insights into the internal representations and feature hierarchies learned by the network for each class for accurate classification.

4. Discussion

High‐resolution quantitative study of features of individual white matter axons in multiple cortical pathways can provide key insights about pathway features, integrity, signal transduction, and neural communication in the neurotypical brain and mechanisms of disruptions in ASD. However, such analyzes require extensive study by experienced anatomists and semiautomated tracing of individual axons. They are thus labor‐intensive and time‐consuming, rendering these approaches not optimal for large‐scale studies that can help identify core ASD network status and likely mechanisms of disruption in communication. A handful of studies have successfully undertaken this task (X. Liu et al. 2020; X. B. Liu and Schumann 2014; Trutzer et al. 2019; Wegiel et al. 2018; Zikopoulos and Barbas 2010; Zikopoulos, Liu, et al. 2018); however, there is a critical need for more detailed studies that will survey relevant cortical and subcortical pathways in neurotypical healthy controls and individuals with ASD. To address this need we developed a multidisciplinary approach that combines sophisticated high‐resolution quantitative microscopy and histopathology with advanced methodologies of machine learning and artificial intelligence (AI). We used our large dataset of light and EM images of white matter axons from short‐ and long‐range cortical pathways below the anterior cingulate cortices in control and ASD groups of adults, from studies published since 2010 (Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018; Zikopoulos, Liu, et al. 2018), to develop and optimize machine learning algorithms. Our results showed that by customizing and optimizing DNNs through transfer learning, we can determine if microscopic images were obtained from postmortem brains of ASD or typically developed individuals, with high confidence (98%). Moreover, we were able to extract image sensitivity maps, to highlight key features in the images that contributed to the accuracy and errors of DNN classification, and evaluate the extent of these contributions, revealing novel insights about mechanisms of ASD pathology.

Overall, the final trained networks demonstrated distinction across four classes (ASD DWM, ASD SWM, CTR DWM, and CTR‐DWM) at multiple scales, spanning cellular (Tol‐Blue) and subcellular (EM) levels. Loading values from principal components revealed crucial details about how the networks learned. The first principal component for both types of images detected the difference between ASD and CTR, which explains the high accuracy of our networks in classifying pathology. The second principal component was primarily focused on white matter region classification. This was reflected in the network's moderately successful performance in classifying white matter regions. Classification across two categories (ASD vs. CTR or SWM vs. DWM), after pooling across a factor (e.g., DWM and SWM) to assess predicted image classes (e.g., ASD vs. CTR) against the true class of each image, bumped up accuracy significantly. Importantly, when the classification results were pooled together across pathology (ASD vs. CTR) and depth (deep vs. SWM), the networks consistently exhibited better performance in classifying pathology compared to white matter regions (Figure 5 and Table 2). Notably, the accuracies for ASD detection in both EM and Tol‐Blue images were very high, but for different white matter regions, when pooled across different white matter depths. Specifically, in EM images, the network predominantly identified all ASD slices as ASD DWM, whereas the Tol‐Blue network primarily classified ASD slices as ASD SWM. This shift in bias in different networks suggests confusion, particularly regarding axons across white matter depths in ASD brains. The difficulty in distinguishing white matter depth led the networks to either predominantly learn traits from ASD DWM and classify these as ASD for EM images or, for Tol‐Blue, learn traits from ASD SWM and generalize those features as ASD. This confusion aligns with known anatomical differences, as the superficial and in some cases the DWM in patients with ASD tend to significantly deviate from the superficial and DWM pattern of neurotypical brains (Ameis and Catani 2015; Barnea‐Goraly et al. 2004; Ecker et al. 2016; Herbert et al. 2004; Herbet et al. 2014; Hong et al. 2019; Jou et al. 2011; X. Liu et al. 2020; X. B. Liu and Schumann 2014; Owen et al. 2014; Radua et al. 2011; Samson et al. 2016; Shen et al. 2024; Shukla, Keehn, and Muller 2011; Shukla, Keehn, Smylie, and Muller 2011; Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018; Zikopoulos, Liu, et al. 2018). Supporting this observation, the loading in the third principal component for ASD SWM, in EM images, was the highest, showing the most dispersion compared to other classes, and indicating that the network identified more variability within ASD SWM (Figure 4). On the other hand, in the second principal component for Tol‐Blue images, the ASD SWM coefficient was the lowest, and an outlier, showing again separation of ASD SWM from the other classes (Figure 4). Both trends, suggest that white matter depth classification accuracy is more challenging in ASD, compared to controls, and establish the central role of SWM changes under ACC in ASD, providing novel insights about a key locus of heterogeneity and pathology in the white matter below ACC.

Importantly, both correct and incorrect classification of images can provide useful insights for the organization of the white matter below ACC in neurotypical individuals and in ASD. The transition between SWM and DWM is gradual and is evident in the coronal sections used here, by the gradual change in the orientation, trajectory variability, and density of axons, as described in previous studies (Zikopoulos and Barbas 2010; Zikopoulos, Garcia‐Cabezas, and Barbas 2018; Zikopoulos, Liu, et al. 2018). Our findings showed that in the ASD group, in particular, the distinction between SWM and DWM was difficult, a novel insight which suggests blending of the white matter regions, due to widespread changes in axon features. Once the DNN learned these features in the SWM or DWM they were then generalizable to all white matter and necessary for ASD classification. Generated image sensitivity maps highlighted specific attributes of images that contributed to DNN classification, providing an opportunity to compare the focus of DNN on each image with key loci and attributes important for neuroanatomical image analyzes. This comparison pinpointed areas of overlap, misses, or novel attributes, converging on the combination of three key features for pathology classification: axon density, heterogeneity of axon sizes, and axon profile shape variability, all of which appeared to be consistently higher in ASD below ACC.

4.1. Limitations, Challenges, and Future Directions

It is well established that axons in the white matter below ACC change in ASD; therefore it is reassuring and not surprising that this machine learning approach could classify the images with high accuracy, after optimization and training of the DNN. A DNN is an artificial neural network with multiple layers between its input and output layers and has simulated neurons, synapses, weights, thresholds, and signal functions. Nonlinear feedforward DNNs with three or more layers are considered “universal approximators,” that is, with the proper architecture and dataset, they can learn to detect and recognize patterns with high accuracy (Hornik et al. 1989). However, the success of a DNN can be significantly constrained by limited training data and limited computational power to perform the training. Moreover, DNNs perform predesigned computations, such as convolutions (hence the interchangeable term convolutional neural network, CNN) that can yield valuable properties such as image position invariance to the classification power, but the presence of such predesigned computations, besides their merits, can limit the DNN domain of universality (Geman et al. 1992). We used multiple methods including DNN transfer learning and augmentation through additional preprocessing of images, fine‐tuning of the number of epochs, and iterations to maximize the accuracy of performance for the images of interest, in line with previous studies (Pan and Yang 2010; Rehman et al. 2020; Ribani and Marengoni 2019; Rustom et al. 2022; Rustom et al. 2024; C. Tan et al. 2018). It must be noted here that differences in the directionality of long‐ and short‐range axons may have played a role in performance biases in network classification across scales. These differences within the SWM are likely minimal, because axons in the SWM follow very variable trajectories, in many cases likely due to fanning out, thinning, and branching as they reach their destinations. In addition, because we consistently used matching coronal sections, we minimized variability across samples and cases. Moreover, even though long‐range axons from the sampled DWM likely continue in the SWM, they only constitute a very small proportion of the total axons (Hilgetag and Zikopoulos 2022; Rosen and Halgren 2022). Nevertheless, even small changes in the pattern and directionality of axons may have an effect on the accuracy of CNN classification. It is likely that the performance bias in Tol‐Blue versus EM classification that we observed, with higher accuracy in the Tol‐Blue versus the EM, was due to two key factors. First, a much larger area of the white matter was included in the samples analyzed for the Tol‐Blue (1.6 mm2/case/region) versus 0.3 mm2/case/region for the EM. Second, EM images were first cropped and then downsampled/compressed to be used as CNN inputs whereas, the Tol‐Blue images were only cropped and were not compressed. Despite these differences, both approaches yielded highly accurate classifications and provided useful insights.

Additionally, we evaluated network performance using dimension reduction techniques, to visualize outliers for further inspection, and using image sensitivity maps, to identify regions and patterns causing DNN confusion or to highlight novel attributes that can be examined. Our findings highlight that even when multiple attributes were combined, to form mixed labels (such as ASD‐DWM) for DNN supervised training, the performance of the network for individual attributes (such as ASD vs. CTR) could be redeemed, paving the way for iterations using a multitude of combined labels, as needed. Our exhaustive training and testing regime with different combinations of nonoverlapping cases (Tables 2 and S1), also revealed that with small datasets, in this case limited number of cases and images, it was best to train separate models for optimal classification of pathology only, or white matter depth only, or combinations of these. Compilation of larger datasets with an increased number of cases and images that can be used for future training and validation can further maximize performance and increase generalizability.

Other recent studies that use machine learning have similarly focused on brain network connectivity, either functional, based on resting‐state functional magnetic resonance imaging (rs‐fMRI) and EEG (Ranaut et al. 2024; Wang et al. 2024), or structural, based on MRI (Aglinskas et al. 2022; Xu et al. 2024) to classify ASD pathology (reviewed in Wen et al. 2024). Our approach is the first to use a multiscale, ultra high‐resolution dataset based on microscopy that can examine and bridge fine structural, molecular, histopathological, and functional pathway attributes to classify patterns of brain network/circuit organization in the neurotypical brain and in ASD. In future studies, we can use the ACC trained network to examine other brain regions that have or do not have known pathology in ASD. Accurate classification of images in a new region “B,” using the already trained ACC‐DNN would suggest that the nature of neuropathology in the two regions (ACC and B) is likely consistent. On the other hand, incorrect classification of images in a new region “B,” using the already trained ACC‐DNN, would suggest either that region “B” is not affected in ASD, or that the nature of neuropathology in the two regions (ACC and B) is different and region dependent. In the latter case, a new, optimized DNN, would need to be developed and trained with images from both ACC and any new region(s) to maximize performance and increase accuracy. Importantly, this approach can be used to examine histopathological white matter samples from a diverse population of individuals with idiopathic and syndromic ASD. The latter includes genetic copy number variations (CNVs) that can lead to developmental delays, intellectual disability and increased risk of ASD, like deletion and duplication of 16p11.2, or 22q13/Shank3 deletion (Phelan‐McDermid syndrome), both of which show extensive axon pathology (Bassell et al. 2020; Marshall et al. 2008; Owen et al. 2014; Pagani et al. 2019; Weiss et al. 2008; Yi et al. 2016; Zhou et al. 2019). This way, our proposed method can be used to parse ASD heterogeneity and identify biomarkers across the spectrum of autism. Moreover, our proposed method can be generalized and applied to detect axonal and pathway feature differences across brain areas in the neurotypical brain and in disorders with underlying pathology in brain network connectivity. Examples include a number of psychiatric and neurologic disorders that include a component of white matter injury or degeneration (e.g., schizophrenia, Alzheimer's, ALS, TBI/CTE, MS; Barbas et al. 2025; Braak and Del Tredici 2018; Graham et al. 2023; Highley et al. 1999; Lewis et al. 2001; Qiu et al. 2010; Shu et al. 2011; Uceda‐Heras et al. 2024; Yao et al. 2020). Notably, our approach at the microscopic scale can be adapted and optimized for the study of the organization of brain pathway structure macroscopically, using in vivo imaging approaches and datasets. This methodical approach holds great promise for the reliable implementation of DNN training and systematic testing to study many brain areas and networks, which would otherwise take prohibitively long to examine, and would provide key insights that will guide future studies and development of targeted diagnostics and interventions.

Author Contributions

Conceptualization: Arash Yazdanbakhsh and Basilis Zikopoulos. Experimental procedures: Basilis Zikopoulos, Xuefeng Liu, and Songlin Xie. Analysis: Kim T.M. Dang, Kelvin Kuang, Tingru Lian, Arash Yazdanbakhsh, and Basilis Zikopoulos. Funding acquisition: Arash Yazdanbakhsh and Basilis Zikopoulos; Writing – original draft: Kim T.M. Dang, Kelvin Kuang, Arash Yazdanbakhsh, and Basilis Zikopoulos. Writing – review and editing: Arash Yazdanbakhsh, Kim T.M. Dang, Kelvin Kuang, Tingru Lian, Xuefeng Liu, Songlin Xie, and Basilis Zikopoulos.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Data S1: Supporting Information.

Acknowledgments

We gratefully acknowledge brain donors and their families, the Harvard Brain Tissue Resource Center, the Institute for Basic Research in Developmental Disabilities, the University of Maryland Brain and Tissue Bank, and the National Disease Research Interchange for providing postmortem human brain tissue. Some tissues were obtained from Autism BrainNet, a resource of the Simons Foundation Autism Research Initiative (SFARI). Autism BrainNet also manages the Autism Tissue Program (ATP) collection, previously funded by Autism Speaks. We are grateful and indebted to the families who donated tissue for research purposes to Autism BrainNet and the ATP. We thank Tara McHugh, for technical assistance in cutting and staining the human specimens. This work was supported by a grant from the Simons Foundation (SFARI Grant 946867) to B.Z. and A.Y.

Yazdanbakhsh, A. , Dang K. T. M., Kuang K., et al. 2025. “Artificial Intelligence Networks Combining Histopathology and Machine Learning Can Extract Axon Pathology in Autism Spectrum Disorder.” Autism Research 18, no. 11: 2210–2230. 10.1002/aur.70135.

Funding: This work was supported by grants from the Simons Foundation Autism Research Initiative (SFARI Grant 946867) and National Institute of Mental Health (R01 MH118500) to B.Z. and A.Y.

Arash Yazdanbakhsh, Kim T. M. Dang, and Kelvin Kuang are co‐first authors.

Contributor Information

Arash Yazdanbakhsh, Email: yazdan@bu.edu.

Basilis Zikopoulos, Email: zikopoul@bu.edu.

Data Availability Statement

The data that support the findings of this study are openly available in Github at https://github.com/kkuang0/ASD‐ACC‐Transfer‐Net.

References

  1. Aglinskas, A. , Hartshorne J. K., and Anzellotti S.. 2022. “Contrastive Machine Learning Reveals the Structure of Neuroanatomical Variation Within Autism.” Science 376, no. 6597: 1070–1074. 10.1126/science.abm2461. [DOI] [PubMed] [Google Scholar]
  2. Ameis, S. H. , and Catani M.. 2015. “Altered White Matter Connectivity as a Neural Substrate for Social Impairment in Autism Spectrum Disorder.” Cortex 62: 158–181. 10.1016/j.cortex.2014.10.014. [DOI] [PubMed] [Google Scholar]
  3. Barbas, H. , Garcia‐Cabezas M. A., John Y., Bautista J., McKee A., and Zikopoulos B.. 2025. “Cortical Circuit Principles Predict Patterns of Trauma Induced Tauopathy in Humans.” Cerebral Cortex 35, no. 8. 10.1093/cercor/bhaf209. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Barnea‐Goraly, N. , Kwon H., Menon V., Eliez S., Lotspeich L., and Reiss A. L.. 2004. “White Matter Structure in Autism: Preliminary Evidence From Diffusion Tensor Imaging.” Biological Psychiatry 55, no. 3: 323–326. [DOI] [PubMed] [Google Scholar]
  5. Bassell, J. , Srivastava S., Prohl A. K., et al. 2020. “Diffusion Tensor Imaging Abnormalities in the Uncinate Fasciculus and Inferior Longitudinal Fasciculus in Phelan‐McDermid Syndrome.” Pediatric Neurology 106: 24–31. 10.1016/j.pediatrneurol.2020.01.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Braak, H. , and Del Tredici K.. 2018. “Spreading of Tau Pathology in Sporadic Alzheimer's Disease Along Cortico‐Cortical Top‐Down Connections.” Cerebral Cortex 28, no. 9: 3372–3384. 10.1093/cercor/bhy152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Burman, P. 1989. “A Comparative Study of Ordinary Cross‐Validation, v‐Fold Cross‐Validation and the Repeated Learning‐Testing Methods.” Biometrika 76, no. 3: 503–514. 10.1093/biomet/76.3.503. [DOI] [Google Scholar]
  8. Caminiti, R. , Carducci F., Piervincenzi C., et al. 2013. “Diameter, Length, Speed, and Conduction Delay of Callosal Axons in Macaque Monkeys and Humans: Comparing Data From Histology and Magnetic Resonance Imaging Diffusion Tractography.” Journal of Neuroscience 33, no. 36: 14501–14511. 10.1523/JNEUROSCI.0761-13.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Chaki, J. , and Deshpande G.. 2024. “Brain Disorder Detection and Diagnosis Using Machine Learning and Deep Learning – A Bibliometric Analysis.” Current Neuropharmacology 22, no. 13: e310524230577. 10.2174/1570159x22999240531160344. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Courchesne, E. , and Pierce K.. 2005. “Why the Frontal Cortex in Autism Might Be Talking Only to Itself: Local Over‐Connectivity But Long‐Distance Disconnection.” Current Opinion in Neurobiology 15, no. 2: 225–230. [DOI] [PubMed] [Google Scholar]
  11. Dichter, G. S. , Felder J. N., and Bodfish J. W.. 2009. “Autism Is Characterized by Dorsal Anterior Cingulate Hyperactivation During Social Target Detection.” Social Cognitive and Affective Neuroscience 4, no. 3: 215–226. http://www.ncbi.nlm.nih.gov/pubmed/19574440. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Ecker, C. , Andrews D., Dell'Acqua F., et al. 2016. “Relationship Between Cortical Gyrification, White Matter Connectivity, and Autism Spectrum Disorder.” Cerebral Cortex 26, no. 7: 3297–3309. 10.1093/cercor/bhw098. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. García‐Cabezas, M. A. , Barbas H., and Zikopoulos B.. 2018. “Parallel Development of Chromatin Patterns, Neuron Morphology, and Connections: Potential for Disruption in Autism.” Frontiers in Neuroanatomy 12: 70. 10.3389/fnana.2018.00070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Geman, S. , Bienenstock E., and Doursat R.. 1992. “Neural Networks and the Bias/Variance Dilemma.” Neural Computation 4, no. 1: 1–58. 10.1162/neco.1992.4.1.1. [DOI] [Google Scholar]
  15. Graham, N. S. N. , Cole J. H., Bourke N. J., Schott J. M., and Sharp D. J.. 2023. “Distinct Patterns of Neurodegeneration After TBI and in Alzheimer's Disease.” Alzheimer's & Dementia 19, no. 7: 3065–3077. 10.1002/alz.12934. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. He, K. , Zhang X., Ren S., and Sun J.. 2015. “Deep Residual Learning for Image Recognition.” http://arxiv.org/abs/1512.03385.
  17. Herbert, M. R. , Ziegler D. A., Makris N., et al. 2004. “Localization of White Matter Volume Increase in Autism and Developmental Language Disorder.” Annals of Neurology 55, no. 4: 530–540. [DOI] [PubMed] [Google Scholar]
  18. Herbet, G. , Lafargue G., Bonnetblanc F., Moritz‐Gasser S., de Menjot Champfleur N., and Duffau H.. 2014. “Inferring a Dual‐Stream Model of Mentalizing From Associative White Matter Fibres Disconnection.” Brain 137, no. Pt 3: 944–959. 10.1093/brain/awt370. [DOI] [PubMed] [Google Scholar]
  19. Highley, J. R. , Esiri M. M., McDonald B., Cortina‐Borja M., Herron B. M., and Crow T. J.. 1999. “The Size and Fibre Composition of the Corpus Callosum With Respect to Gender and Schizophrenia: A Post‐Mortem Study.” Brain 122, no. Pt 1: 99–110. [DOI] [PubMed] [Google Scholar]
  20. Hilgetag, C. C. , and Zikopoulos B.. 2022. “The Highways and Byways of the Brain.” PLoS Biology 20, no. 3: e3001612. 10.1371/journal.pbio.3001612. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Hill, E. L. 2004. “Executive Dysfunction in Autism.” Trends in Cognitive Sciences 8, no. 1: 26–32. [DOI] [PubMed] [Google Scholar]
  22. Hong, S. J. , Hyung B., Paquola C., and Bernhardt B. C.. 2019. “The Superficial White Matter in Autism and Its Role in Connectivity Anomalies and Symptom Severity.” Cerebral Cortex 29, no. 10: 4415–4425. 10.1093/cercor/bhy321. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Hornik, K. , Stinchcombe M., and White H.. 1989. “Multilayer Feedforward Networks Are Universal Approximators.” Neural Networks 2, no. 5: 359–366. 10.1016/0893-6080(89)90020-8. [DOI] [Google Scholar]
  24. Iandola, F. N. , Han S., Moskewicz M. W., Ashraf K., Dally W. J., and Keutzer K.. 2016. “SqueezeNet: AlexNet‐Level Accuracy With 50x Fewer Parameters and <0.5MB Model Size.” http://arxiv.org/abs/1602.07360.
  25. Innocenti, G. M. , and Caminiti R.. 2017. “Axon Diameter Relates to Synaptic Bouton Size: Structural Properties Define Computationally Different Types of Cortical Connections in Primates.” Brain Structure & Function 222, no. 3: 1169–1177. 10.1007/s00429-016-1266-1. [DOI] [PubMed] [Google Scholar]
  26. Jou, R. J. , Mateljevic N., Minshew N. J., Keshavan M. S., and Hardan A. Y.. 2011. “Reduced Central White Matter Volume in Autism: Implications for Long‐Range Connectivity.” Psychiatry and Clinical Neurosciences 65, no. 1: 98–101. 10.1111/j.1440-1819.2010.02164.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Kana, R. K. , Keller T. A., Minshew N. J., and Just M. A.. 2006. “Inhibitory Control in High‐Functioning Autism: Decreased Activation and Underconnectivity in Inhibition Networks.” Biological Psychiatry 62, no. 3: 198–206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Krizhevsky, A. , Sutskever I., and Hinton G. E.. 2012. ImageNet Classification With Deep Convolutional Neural Networks, edited by Pereira F., Burges C. J., Bottou L., and Weinberger K. Q., vol. 25. Curran Associates, Inc. [Google Scholar]
  29. Kundu, S. , Sair H., Sherr E. H., Mukherjee P., and Rohde G. K.. 2024. “Discovering the Gene‐Brain‐Behavior Link in Autism via Generative Machine Learning.” Science Advances 10, no. 24: eadl5307. 10.1126/sciadv.adl5307. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Lewis, D. A. , Cruz D. A., Melchitzky D. S., and Pierri J. N.. 2001. “Lamina‐Specific Deficits in Parvalbumin‐Immunoreactive Varicosities in the Prefrontal Cortex of Subjects With Schizophrenia: Evidence for Fewer Projections From the Thalamus.” American Journal of Psychiatry 158, no. 9: 1411–1422. [DOI] [PubMed] [Google Scholar]
  31. Liu, X. , Bautista J., Liu E., and Zikopoulos B.. 2020. “Imbalance of Laminar‐Specific Excitatory and Inhibitory Circuits of the Orbitofrontal Cortex in Autism.” Molecular Autism 11, no. 1: 83. 10.1186/s13229-020-00390-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Liu, X. B. , and Schumann C. M.. 2014. “Optimization of Electron Microscopy for Human Brains With Long‐Term Fixation and Fixed‐Frozen Sections.” Acta Neuropathologica Communications 2: 42. 10.1186/2051-5960-2-42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Mai, J. K. , Majtanik M., and Paxinos G.. 2015. Atlas of the Human Brain. 4th ed. Academic Press – Elsevier. [Google Scholar]
  34. Marshall, C. R. , Noor A., Vincent J. B., et al. 2008. “Structural Variation of Chromosomes in Autism Spectrum Disorder.” American Journal of Human Genetics 82, no. 2: 477–488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Mordvintsev, A. , Olah C., and Tyka M.. 2015. “DeepDream – A Code Example for Visualizing Neural Networks.” https://research.google/blog/deepdream‐a‐code‐example‐for‐visualizing‐neural‐networks/.
  36. Owen, J. P. , Chang Y. S., Pojman N. J., et al. 2014. “Aberrant White Matter Microstructure in Children With 16p11.2 Deletions.” Journal of Neuroscience 34, no. 18: 6214–6223. 10.1523/JNEUROSCI.4495-13.2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Pagani, M. , Bertero A., Liska A., et al. 2019. “Deletion of Autism Risk Gene Shank3 Disrupts Prefrontal Connectivity.” Journal of Neuroscience 39, no. 27: 5299–5310. 10.1523/jneurosci.2529-18.2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Palomero‐Gallagher, N. , Hoffstaedter F., Mohlberg H., Eickhoff S. B., Amunts K., and Zilles K.. 2018. “Human Pregenual Anterior Cingulate Cortex: Structural, Functional, and Connectional Heterogeneity.” Cerebral Cortex 29: 2552–2574. 10.1093/cercor/bhy124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Palomero‐Gallagher, N. , Mohlberg H., Zilles K., and Vogt B.. 2008. “Cytology and Receptor Architecture of Human Anterior Cingulate Cortex.” Journal of Comparative Neurology 508, no. 6: 906–926. 10.1002/cne.21684. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Palomero‐Gallagher, N. , Zilles K., Schleicher A., and Vogt B. A.. 2013. “Cyto‐ and Receptor Architecture of Area 32 in Human and Macaque Brains.” Journal of Comparative Neurology 521, no. 14: 3272–3286. 10.1002/cne.23346. [DOI] [PubMed] [Google Scholar]
  41. Pan, S. J. , and Yang Q.. 2010. “A Survey on Transfer Learning.” IEEE Transactions on Knowledge and Data Engineering 22, no. 10: 1345–1359. 10.1109/tkde.2009.191. [DOI] [Google Scholar]
  42. Pedregosa, F. , Varoquaux G., Gramfort A., et al. 2011. “Scikit‐Learn: Machine Learning in Python.” Journal of Machine Learning Research 12: 2825–2830. [Google Scholar]
  43. Powers, D. 2008. “Evaluation: From Precision, Recall and F‐Factor to ROC, Informedness, Markedness & Correlation.” Machine Learning Technologies 2: 37–63. [Google Scholar]
  44. Qiu, A. , Tuan T. A., Woon P. S., Abdul‐Rahman M. F., Graham S., and Sim K.. 2010. “Hippocampal‐Cortical Structural Connectivity Disruptions in Schizophrenia: An Integrated Perspective From Hippocampal Shape, Cortical Thickness, and Integrity of White Matter Bundles.” NeuroImage 52, no. 4: 1181–1189. [DOI] [PubMed] [Google Scholar]
  45. Radua, J. , Via E., Catani M., and Mataix‐Cols D.. 2011. “Voxel‐Based Meta‐Analysis of Regional White‐Matter Volume Differences in Autism Spectrum Disorder Versus Healthy Controls.” Psychological Medicine 41, no. 7: 1539–1550. [DOI] [PubMed] [Google Scholar]
  46. Ranaut, A. , Khandnor P., and Chand T.. 2024. “Identifying Autism Using EEG: Unleashing the Power of Feature Selection and Machine Learning.” Biomedical Physics & Engineering Express 10, no. 3: 035013. 10.1088/2057-1976/ad31fb. [DOI] [PubMed] [Google Scholar]
  47. Rehman, A. , Naz S., Razzak M. I., Akram F., and Imran M.. 2020. “A Deep Learning‐Based Framework for Automatic Brain Tumors Classification Using Transfer Learning.” Circuits, Systems, and Signal Processing 39, no. 2: 757–775. 10.1007/s00034-019-01246-3. [DOI] [Google Scholar]
  48. Ribani, R. , and Marengoni M.. 2019. “A Survey of Transfer Learning for Convolutional Neural Networks.” Paper presented at the 2019 32nd SIBGRAPI Conference on Graphics, Patterns and Images Tutorials (SIBGRAPI‐T).
  49. Rosen, B. Q. , and Halgren E.. 2022. “An Estimation of the Absolute Number of Axons Indicates That Human Cortical Areas Are Sparsely Connected.” PLoS Biology 20, no. 3: e3001575. 10.1371/journal.pbio.3001575. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Rustom, F. , Ogmen H., and Yazdanbahksh A.. 2022. “Object Detection, Recognition, Deep Learning, and the Universal Law of Generalization.” arXiv. 10.48550/arXiv.2206.05365. [DOI]
  51. Rustom, F. , Parva P., Ogmen H., and Yazdanbakhsh A.. 2024. “Deep Learning and Transfer Learning for Brain Tumor Detection and Classification.” bioRxiv. 10.1101/2023.04.10.536226. [DOI] [PMC free article] [PubMed]
  52. Samson, A. C. , Dougherty R. F., Lee I. A., Phillips J. M., Gross J. J., and Hardan A. Y.. 2016. “White Matter Structure in the Uncinate Fasciculus: Implications for Socio‐Affective Deficits in Autism Spectrum Disorder.” Psychiatry Research: Neuroimaging 255: 66–74. 10.1016/j.pscychresns.2016.08.004. [DOI] [PubMed] [Google Scholar]
  53. Schmahmann, J. D. , and Pandya D. N.. 2006. Fiber Pathways of the Brain. Oxford University Press Inc. [Google Scholar]
  54. Shen, Y. , Zhao X., Wang K., et al. 2024. “Exploring White Matter Abnormalities in Young Children With Autism Spectrum Disorder: Integrating Multi‐Shell Diffusion Data and Machine Learning Analysis.” Academic Radiology 31, no. 5: 2074–2084. 10.1016/j.acra.2023.12.023. [DOI] [PubMed] [Google Scholar]
  55. Shu, N. , Liu Y., Li K. C., et al. 2011. “Diffusion Tensor Tractography Reveals Disrupted Topological Efficiency in White Matter Structural Networks in Multiple Sclerosis.” Cerebral Cortex 21, no. 11: 2565–2577. [DOI] [PubMed] [Google Scholar]
  56. Shukla, D. K. , Keehn B., and Muller R. A.. 2011. “Tract‐Specific Analyses of Diffusion Tensor Imaging Show Widespread White Matter Compromise in Autism Spectrum Disorder.” Journal of Child Psychology and Psychiatry 52, no. 3: 286–295. 10.1111/j.1469-7610.2010.02342.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Shukla, D. K. , Keehn B., Smylie D. M., and Muller R. A.. 2011. “Microstructural Abnormalities of Short‐Distance White Matter Tracts in Autism Spectrum Disorder.” Neuropsychologia 49, no. 5: 1378–1382. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Simonyan, K. , and Zisserman A.. 2014. “Very Deep Convolutional Networks for Large‐Scale Image Recognition.” http://arxiv.org/abs/1409.1556.
  59. Szegedy, C. , Liu W., Jia Y., et al. 2014. “Going Deeper With Convolutions.” http://arxiv.org/abs/1409.4842.
  60. Szegedy, C. , Wei L., Yangqing J., et al. 2015. “Going Deeper With Convolutions.” Paper presented at the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  61. Tan, C. , Sun F., Kong T., Zhang W., Yang C., and Liu C.. 2018. “A Survey on Deep Transfer Learning.” arXiv. 10.48550/arXiv.1808.01974. [DOI]
  62. Tan, M. , and Le Q. V.. 2019. “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.” http://arxiv.org/abs/1905.11946.
  63. Thakkar, K. N. , Polli F. E., Joseph R. M., et al. 2008. “Response Monitoring, Repetitive Behaviour and Anterior Cingulate Abnormalities in Autism Spectrum Disorders (ASD).” Brain 131, no. Pt 9: 2464–2478. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Trutzer, I. M. , Garcia‐Cabezas M. A., and Zikopoulos B.. 2019. “Postnatal Development and Maturation of Layer 1 in the Lateral Prefrontal Cortex and Its Disruption in Autism.” Acta Neuropathologica Communications 7, no. 1: 40. 10.1186/s40478-019-0684-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Uceda‐Heras, A. , Aparicio‐Rodriguez G., and Garcia‐Cabezas M. A.. 2024. “Hyperphosphorylated Tau in Alzheimer's Disease Disseminates Along Pathways Predicted by the Structural Model for Cortico‐Cortical Connections.” Journal of Comparative Neurology 532, no. 5: e25623. 10.1002/cne.25623. [DOI] [PubMed] [Google Scholar]
  66. Vogt, B. A. , Hof P. R., Zilles K., Vogt L. J., Herold C., and Palomero‐Gallagher N.. 2013. “Cingulate Area 32 Homologies in Mouse, Rat, Macaque and Human: Cytoarchitecture and Receptor Architecture.” Journal of Comparative Neurology 521, no. 18: 4189–4204. 10.1002/cne.23409. [DOI] [PubMed] [Google Scholar]
  67. von Economo, C. , and Koskinas G. N.. 1925/2008. Atlas of Cytoarchitectonics of the Adult Human Cerebral Cortex. Translated From the German Original, Revised and Edited With an Introduction and Additional Appendix Material by L. C. Triarhou. 1st English ed. Karger. [Google Scholar]
  68. Wang, X. , Zhao K., Yao L., et al. 2024. “Delineating Transdiagnostic Subtypes in Neurodevelopmental Disorders via Contrastive Graph Machine Learning of Brain Connectivity Patterns.” bioRxiv. 10.1101/2024.02.29.582790. [DOI]
  69. Wegiel, J. , Kaczmarski W., Flory M., et al. 2018. “Deficit of Corpus Callosum Axons, Reduced Axon Diameter and Decreased Area Are Markers of Abnormal Development of Interhemispheric Connections in Autistic Subjects.” Acta Neuropathologica Communications 6, no. 1: 143. 10.1186/s40478-018-0645-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Weiss, L. A. , Shen Y., Korn J. M., et al. 2008. “Association Between Microdeletion and Microduplication at 16p11.2 and Autism.” New England Journal of Medicine 358, no. 7: 667–675. [DOI] [PubMed] [Google Scholar]
  71. Wen, J. , Antoniades M., Yang Z., et al. 2024. “Dimensional Neuroimaging Endophenotypes: Neurobiological Representations of Disease Heterogeneity Through Machine Learning.” Biological Psychiatry 96: 564–584. 10.1016/j.biopsych.2024.04.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Xu, G. , Geng G., Wang A., et al. 2024. “Three Autism Subtypes Based on Single‐Subject Gray Matter Network Revealed by Semi‐Supervised Machine Learning.” Autism Research 17: 1962–1973. 10.1002/aur.3183. [DOI] [PubMed] [Google Scholar]
  73. Yadav, S. S. , and Jadhav S. M.. 2019. “Deep Convolutional Neural Network Based Medical Image Classification for Disease Diagnosis.” Journal of Big Data 6, no. 1: 113. 10.1186/s40537-019-0276-2. [DOI] [Google Scholar]
  74. Yao, B. , Neggers S. F. W., Kahn R. S., and Thakkar K. N.. 2020. “Altered Thalamocortical Structural Connectivity in Persons With Schizophrenia and Healthy Siblings.” Neuroimage: Clinical 28: 102370. 10.1016/j.nicl.2020.102370. [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Yi, F. , Danko T., Botelho S. C., et al. 2016. “Autism‐Associated SHANK3 Haploinsufficiency Causes Ih Channelopathy in Human Neurons.” Science 352, no. 6286: aaf2669. 10.1126/science.aaf2669. [DOI] [PMC free article] [PubMed] [Google Scholar]
  76. Zhou, Y. , Sharma J., Ke Q., et al. 2019. “Atypical Behaviour and Connectivity in SHANK3‐Mutant Macaques.” Nature 570, no. 7761: 326–331. 10.1038/s41586-019-1278-0. [DOI] [PubMed] [Google Scholar]
  77. Zikopoulos, B. , and Barbas H.. 2010. “Changes in Prefrontal Axons May Disrupt the Network in Autism.” Journal of Neuroscience 30, no. 44: 14595–14609. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. Zikopoulos, B. , and Barbas H.. 2013. “Altered Neural Connectivity in Excitatory and Inhibitory Cortical Circuits in Autism.” Frontiers in Human Neuroscience 7: 609. 10.3389/fnhum.2013.00609. [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. Zikopoulos, B. , Garcia‐Cabezas M. A., and Barbas H.. 2018. “Parallel Trends in Cortical Grey and White Matter Architecture and Connections in Primates Allow Fine Study of Pathways in Humans and Reveal Network Disruptions in Autism.” PLoS Biology 16, no. 2: e2004559. 10.1371/journal.pbio.2004559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  80. Zikopoulos, B. , Liu X., Tepe J., Trutzer I., John Y. J., and Barbas H.. 2018. “Opposite Development of Short‐ and Long‐Range Anterior Cingulate Pathways in Autism.” Acta Neuropathologica 136, no. 5: 759–778. 10.1007/s00401-018-1904-1. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1: Supporting Information.

Data Availability Statement

The data that support the findings of this study are openly available in Github at https://github.com/kkuang0/ASD‐ACC‐Transfer‐Net.


Articles from Autism Research are provided here courtesy of Wiley

RESOURCES