Skip to main content
Biological Psychiatry Global Open Science logoLink to Biological Psychiatry Global Open Science
. 2026 Apr 23;6(4):100745. doi: 10.1016/j.bpsgos.2026.100745

Neuroimaging-Based Subgroups in Schizophrenia: A Critical Appraisal of Clustering Studies

Yuetong Yu a, Ruiyang Ge a, Sophia Frangou a,b,∗
PMCID: PMC13272521  PMID: 42318040

Abstract

Efforts to define biologically grounded subtypes of schizophrenia have increasingly leveraged neuroimaging data and clustering algorithms. Such approaches aim to capture patient-level heterogeneity with potential clinical and mechanistic relevance. In this review, we evaluated whether subtypes derived solely from neuroimaging data can be robustly identified and meaningfully linked to clinical variation. A systematic review of peer-reviewed studies published between January 2015 and December 2024 that applied data-driven clustering algorithms to neuroimaging data was conducted to identify patient-level subtypes in individuals with schizophrenia or related spectrum disorders. We excluded transdiagnostic studies, studies focused solely on case-control classification, studies that included variables beyond neuroimaging measures (e.g., clinical or cognitive features) in the clustering input, and studies that performed feature-level clustering without assigning individual-level subtypes. Eighteen studies met inclusion criteria. These studies used structural magnetic resonance imaging features as input. Both the features and the clustering algorithms used varied widely. Across studies, 3 broad neuroanatomical patterns were identified: subtypes with widespread abnormalities, those with regionally circumscribed abnormalities, and those with largely preserved profiles. However, the specific brain regions implicated within each subtype varied considerably between studies, and no subtype profile was consistently reproduced. Few studies reported associations between subtypes and clinical features. When such associations were detected, subtypes characterized by more widespread structural abnormalities tended to show higher symptom severity. Current evidence is insufficient to determine whether macroscale neuroimaging features can define subtypes of schizophrenia that are reproducible, biologically valid, and clinically meaningful. The subtypes reported to date may instead reflect continuous variation within the disorder rather than discrete, biologically distinct entities. Advancing the field will require larger, harmonized datasets, standardized analytic pipelines, and rigorous external and longitudinal validation.

Keywords: Clustering, Machine learning, Neuroimaging, Schizophrenia, Subgroup, Subtyping

Plain Language Summary

Schizophrenia is a heterogeneous disorder, and multiple studies have attempted to identify brain-based subtypes of the disorder using clustering methods applied to neuroimaging data. In this review, we examined studies published between 2015 and 2024 that used neuroimaging features alone to assign individual patients to brain-based subtypes.

We identified 18 eligible studies that used a diverse array of MRI measures and unsupervised or semisupervised clustering approaches. Across studies, 3 broad types of subgroups were commonly reported: patients with widespread brain abnormalities, patients with more localized abnormalities, and patients with largely preserved brain structure. However, the number of subtypes and the specific brain regions involved varied substantially between studies. Subtype differences were rarely linked to clinical features. When associations were reported, patients with more widespread brain abnormalities tended to show greater symptom severity or poorer functioning.

Overall, the findings suggest that although neuroimaging can capture brain-based heterogeneity in schizophrenia, current methods do not yet yield consistent or clinically meaningful subtypes. Larger samples, improved methodological consistency, and stronger validation are needed before neuroimaging-based subtyping can be reliably applied.

Plain Language Summary

Schizophrenia is a heterogeneous disorder, and multiple studies have attempted to identify brain-based subtypes of the disorder using clustering methods applied to neuroimaging data. In this review, we examined studies published between 2015 and 2024 that used neuroimaging features alone to assign individual patients to brain-based subtypes.

We identified 18 eligible studies that used a diverse array of MRI measures and unsupervised or semisupervised clustering approaches. Across studies, 3 broad types of subgroups were commonly reported: patients with widespread brain abnormalities, patients with more localized abnormalities, and patients with largely preserved brain structure. However, the number of subtypes and the specific brain regions involved varied substantially between studies. Subtype differences were rarely linked to clinical features. When associations were reported, patients with more widespread brain abnormalities tended to show greater symptom severity or poorer functioning.

Overall, the findings suggest that although neuroimaging can capture brain-based heterogeneity in schizophrenia, current methods do not yet yield consistent or clinically meaningful subtypes. Larger samples, improved methodological consistency, and stronger validation are needed before neuroimaging-based subtyping can be reliably applied.


Biological subtyping seeks to replace symptom-based taxonomies with classifications grounded in measurable biological features. In the computational literature, subtyping is used in 2 distinct ways. One approach describes heterogeneity by extracting continuous patterns of variation across individuals using unsupervised decomposition methods, such as principal component analysis or non-negative matrix factorization. These methods summarize shared patterns of variation at the group level, but they do not support individual-level classification of patients into distinct subtypes. A second approach uses unsupervised clustering to group individual patients into subtypes based on the similarity of their biological profiles. This is typically implemented with clustering algorithms such as k-means and hierarchical agglomerative clustering, which partition individuals into discrete groups based on their similarity, or with probabilistic mixture models, such as Gaussian mixture models, which explicitly model the data as arising from a mixture of latent subpopulations. The latter approach, which supports biologically informed individual-level classification, is central to the paradigm of precision medicine. Cancer care leads the field, with molecular subtypes already embedded in diagnostic and treatment protocols (1,2). A comparable emphasis on individual-level classification is increasingly seen in neurology, where computational modeling of neuroimaging data has been used to derive biologically distinct subgroups in conditions such as multiple sclerosis (3) and Parkinson’s disease (4). This individual-level orientation aligns with clinical practice and underpins the success of biological classifications used in medicine.

Among psychiatric disorders, schizophrenia has long been regarded as a prototypically heterogeneous syndrome characterized by interindividual variability in positive and negative symptoms, cognitive deficits, and longitudinal course (5,6). Neuroimaging has played a pivotal role in establishing schizophrenia as a brain disorder by revealing reproducible alterations in brain structure, function, and connectivity (7, 8, 9). Neuroimaging has also been the most extensively used modality in efforts to identify biological subtypes of schizophrenia (10) largely because its measures are considered relatively stable and more proximally related to underlying neurobiological mechanisms than clinical or behavioral assessments. The increasing availability of large-scale neuroimaging datasets coupled with advances in standardized preprocessing pipelines has further established neuroimaging as a central data source for biological subtyping. A wide range of clustering approaches has been applied to neuroimaging data in an effort to delineate subtypes of individuals with schizophrenia based on shared brain-based characteristics that may support clinical stratification or point toward distinct pathophysiological or etiological mechanisms within the broader diagnosis.

This review synthesizes findings from clustering studies in schizophrenia in which only algorithms that assign each patient to a brain-based subgroup were used. We further restricted clustering inputs to structural neuroimaging measures to test whether brain features alone can yield reproducible patient subgroups. Including clinical or other biological variables would define multimodal clusters, making it difficult to attribute subgroup structure specifically to neuroimaging. We reviewed studies published between January 1, 2015, and December 31, 2024, to capture the most recent and methodologically advanced applications and evaluate whether neuroimaging-derived subtypes can be robustly identified and meaningfully linked to clinical variation within schizophrenia. Our appraisal highlights methodological strengths, limitations, and reproducibility issues to inform future research on neuroimaging-grounded patient stratification.

Methods and Materials

This review evaluated original peer-reviewed studies published in English in the past 10 years (January 1, 2015 to December 31, 2024) and referenced in the major databases (see the Supplement for details). Inclusion criteria included the following: 1) studies that applied clustering algorithms to group individuals with schizophrenia and spectrum disorders into discrete or probabilistic subtypes, 2) studies that used only neuroimaging features as input to the clustering algorithm, and 3) studies that explicitly reported the number and defining characteristics of the resulting subtypes.

Exclusion criteria included the following:

  • 1.

    Studies using clustering to identify population-level patterns of neuroimaging features without deriving participant-level subtype assignments.

  • 2.

    Studies that incorporated additional modalities (e.g., clinical, cognitive, or genetic data) as input features to the clustering algorithms. This restriction was imposed to allow an evaluation of the capacity of neuroimaging features to capture biologically meaningful heterogeneity within schizophrenia; multidomain input features would have made it difficult to isolate the specific contribution of neuroimaging to subtype differentiation.

  • 3.

    Studies that focused exclusively on classification (i.e., distinguishing individuals with schizophrenia from healthy control participants) without deriving within-group subtypes.

  • 4.

    Studies that applied clustering to mixed samples comprising patients with schizophrenia and either healthy control participants or unaffected relatives, as such designs risk deriving clusters that primarily reflect diagnostic group differences rather than meaningful heterogeneity within the patient population.

  • 5.

    Studies that applied clustering to mixed clinical samples comprising patients with schizophrenia together with patients with other psychiatric diagnoses (e.g., mood disorders). This decision was made to avoid conflating transdiagnostic variation with within-diagnosis heterogeneity. In mixed-diagnosis samples, clustering algorithms are likely to capture differences driven by diagnostic differences or similarities, thereby obscuring the disorder-specific neuroanatomical patterns necessary for identifying meaningful subtypes within schizophrenia. This problem is exacerbated when diagnostic groups in mixed samples differ in age, sex, sample size, illness stage, or medication exposure, as these imbalances can bias clustering toward features linked to such confounders rather than to the intrinsic heterogeneity of schizophrenia.

  • 6.

    Preprints, reviews, meta-analyses, commentaries, letters, and other nonprimary research articles were also excluded.

Titles and abstracts were screened independently by 2 of the authors to determine eligibility, followed by full-text review of all potentially relevant articles. Disagreements were resolved through discussion and consensus.

Results

We identified 18 studies that met the selection criteria. The Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) checklist and flowchart are included in the Supplement. Details of the individual studies and their main findings are presented in Table 1, and information about study quality and validation practices is provided in Table 2. Below we provide a synthesis of the key observations and findings organized thematically.

Table 1.

Summary of Studies

Study Samplea Clinical Stageb Input Features for Clustering Clustering Method(s) Number of Clusters Subtype Descriptionc Associated Featuresd
Volumetric Measures

Gupta et al., 2017 (24) 382 patients ∼72% male; ∼avg. age: 36 years Established ICA-derived patterns of voxelwise GMV Biclustered independent component analysis 2 Subtype 1 (regional): reduced gray matter concentration in the insula, superior temporal gyrus, and inferior frontal gyrus
Subtype 2 (regional): reduced gray matter concentration in the superior, middle, and medial frontal gyri; 202 patients remained unassigned
No subtype differences in age or sex; subtype 2 had higher PANSS positive scores than subtype 1, with no differences in negative or general scores.
Dwyer et al., 2018 (16) 145 (71 patients and 74 HCs); ∼74% male; ∼avg. age: 38 years Established PCA-derived patterns of voxelwise GMV Fuzzy c-means
Clustering after PCA
2 Subtype 1 (widespread): GMV reductions in the insula, striatum, thalamus, hippocampus, and right superior temporal regions, with increased volume in medial and lateral parietal lobes
Subtype 2 (widespread): GMV reductions in lateral prefrontal, medial parietal, and temporal cortices, with increased cerebellar volume
Subtype 1 had an older age of onset, longer duration of illness, and higher PANSS negative symptom scores; no differences were noted in antipsychotic dose/type.
Honnorat et al., 2019 (30) 326 (157 patients and 169 HCs); ∼71% male; ∼avg. age: 31 years Early stage ROI-based GMV, WMV, and CSF volumes CHIMERA 3 Subtype 1 (widespread): reduced GMV in the thalamus, anterior cingulate, and superior temporal regions and higher CSF volume in temporal regions, WM expansion in temporal regions
Subtype 2 (widespread): similar GMV pattern to subtype1 but increased CSF volume in prefrontal regions
Subtype 3 (Preserved): Mild CSF expansion and moderate GM and white matter volume reductions.
Subtype 1 had older age and subtype 3 included more females; no subtype differences were noted in illness duration, age of onset, PANSS subscale scores and antipsychotic dose.
Ma et al., 2019 (21) 67 patients
Unspecified age and sex
FEP drug-naïve ROI-based GMV K-means 3 Subtype 1 (regional): reduced GMV in the supramarginal gyrus and precuneus and higher GMV in the mid-cingulate cortex
Subtype 2 (regional): increased GMV in the angular gyrus
Subtype 3 (preserved): Subtle GMV reduction of the precuneus
No subtype differences in age and sex; subtype 1 had higher PANSS positive symptom scores.
Chand et al., 2020 (19) 671 (307 patients and 364 HCs); ∼60% male; ∼avg. age: 30 years Established ROI-based GMV, WMV, and CSF volumes HYDRA 2 Subtype 1 (Widespread): Widespread volumetric decreases
Subtype 2 (Preserved): Preserved neuroanatomical profile with subtle enlargement of the basal ganglia.
No subtype differences were noted in age, sex, illness duration, antipsychotic dose/type, age of illness onset, or symptom severity.
Liu et al., 2021 (17) 178 (107 patients and 71 HCs); ∼60% male; ∼avg. age: 25 years FEP drug-naïve Structural covariance network constructed by ROI-based GMV K-means 2 Subtype 1 (preserved): Preserved covariance profile
Subtype 2 (regional): decreased covariance between the hippocampus/pallidum/putamen and between occipital/orbitofrontal/superior temporal cortical regions.
No subtype differences in age, sex, illness duration, or PANSS positive/negative symptom scores; the same pattern was identified in 2 FEP datasets, 2 datasets with chronic patients, and in a clinical high-risk sample.
Chew et al., 2022 (13) 234 (158 patients and 76 HCs); ∼68% male; ∼avg. age: 32 years Established ROI-based cortical, subcortical, cerebellar, brainstem, and ventricle volumes HYDRA 2 Subtype 1 (regional): regional increase in the basal ganglia and third ventricle
Subtype 2 (widespread): Widespread cortical and subcortical volume decrease
No subtype differences in age, sex, PANSS total score, or GAF scores; subtype 1 had a longer duration of illness.
Dwyer et al., 2023 (18) 996 (572 patients and 424 HCs); ∼67% male; ∼avg. age: 26 years FEP ROI-based GMV and WMV HYDRA 4 Half of the individuals with FEP were assigned to the previously defined subtype 1 (32%) (widespread) and subtype 2 (21%) (preserved); a small percentage (9%) of individuals with FEP were assigned to subtype 3 (widespread), which had mixed subtype 1 and 2 features, and the remainder were not assigned to any subtype. No subtype differences in age and sex; subtype 1 had lower PANSS scores and included a greater proportion of individuals prescribed typical antipsychotics.
Jiang et al., 2023 (20) 2170 (1124 patients and 1046 HCs); ∼55% male; ∼avg. age: 32 years Mixed: established and FEP ROI-based GMV SuStaIn 2 Both subtypes involved spatially overlapping GMV loss across cortical and subcortical regions. In subtype 1 (widespread), early GMV reductions were mostly prefrontal, and in subtype 2 (widespread) they were mostly subcortical. Across both subtypes, PANSS negative scores were higher with advancing gray matter loss, and the opposite was true with PANSS positive symptom scores.
Jiang et al., 2024 (11) 11,260 (4222 patients and 7038 HCs); ∼55% male; ∼avg age: 33 years Mixed: FEP but mostly established ROI-based GMV SuStaIn 2 Both subtypes showed reductions in cortical thickness and volume and subcortical volume; the magnitude of cortical reductions was greater in frontal regions for subtype 1 (widespread) and in temporal regions for subtype 2 (widespread); the magnitude of subcortical reductions was greater in subtype 2 than in subtype 1 except for the striatum, which was enlarged in subtype 1 and reduced in subtype 2. No subtype differences in age, sex, illness duration, or any of the PANSS subscale scores.
Chai et al., 2023 (12) 571 (314 patients and 257 HCs); ∼69% male; ∼avg. age: 34 years Mixed: established and drug- naïve patients with FEP PCA-derived patterns of voxelwise GMV K-means and agglomerative hierarchical clustering 2 Subtype 1 (widespread): widespread GMV decrease
Subtype 2 (regional): Focally increased GMV in prefrontal and premotor areas; both clustering methods yielded similar subtypes.
No subtype differences in age, sex, or any PANSS item scores; subtype 1 had lower age of onset and lower antipsychotic dose.
Wen et al., 2022 (29) 1166 (583 patients and 583 HCs); ∼62% male; ∼avg. age: 33 years Established opNMF-derived loading coefficients of voxelwise GMV HYDRA 2–4 Plausible solutions identified 2–4 clusters; the 2-cluster solution was chosen because these subtypes were similar to those reported by Chand et al. (19).
Subtype 1 (widespread): widespread gray matter reduction, most pronounced in the insula and thalamus
Subtype 2 (preserved): preserved neuroanatomical profile with subtle enlargement of the basal ganglia
No subtype differences in age, sex, PANSS subscale scores, illness duration, or age of onset; subtype 2 had a higher average GAF score.
Yakimov et al., 2024 (23) 239 (132 patients and 107 HC); ∼71% male; ∼avg age: 37 years Established Choroid plexus, ventricles volumes K-means 3 Subtype 1 (regional): increased choroid plexus and ventricular volumes
Subtype 2 (preserved): lower choroid plexus and ventricular volumes
Subtype 3 (regional): moderately increased choroid plexus and lower ventricular volumes
No subtype differences in age, sex, duration of untreated psychosis or duration of illness, PANSS subscale scores, proportion of treatment-resistant cases, body mass index, or antipsychotic dose; all subtypes showed cognitive impairment, which was greater for subtype 1.

Surface-Based and Volumetric Measures

Yao et al., 2023 (28) 206 (100 patients and 106 HCs); ∼66% male; ∼avg. age: 14 years Childhood and adolescent onset psychosis Morphometric similarity networks derived from measures of surface area, cortical thickness, cortical volume, Gaussian curvature, and mean curvature HYDRA 2 Subtype 1 (widespread): reduction in global MSN strength and in frontotemporal and insula regions coupled with decreased MSN strength in ventral prefrontal and occipital regions
Subtype 2 (regional): normal global MSN strength, increased MSN strength in the superior frontal cortex, and decreased MSN in the paracentral cortex
No subtype differences in age, sex, or PANSS subscale scores
Xiao et al., 2022 (15) FEP: 336 (163 patients and 173 HCs); ∼49% male; ∼avg. age: 24 years
Established: 363 (133 patients and 230 HCs); ∼53% male; ∼avg. age: 33 years
FEP drug-naïve and established PCA-derived patterns of ROI-based cortical thickness, surface area, and subcortical volume Density peak-based clustering 3 FEP: subtype 1 (widespread) showed a widespread decrease in cortical thickness coupled with a focal increase in the rostral anterior cingulate
Subtypes 2 (preserved) and 3 (preserved) had preserved neuroanatomical profiles
Established: subtype 1 (widespread) showed widespread cortical deficits coupled with reduced hippocampal volume and increased pallidal volume
Subtype 2 (regional) had a focal decrease in hippocampal volume
Subtype 3 (preserved) had a preserved neuroanatomical profile
Within each dataset, subtypes did not differ in age, sex (except for a slight excess of females in FEP subtype 2), PANSS subscale scores, GAF, or antipsychotic dose.
Sone et al., 2024 (32) 250 (177 patients and 73 HCs); ∼59% male; ∼avg. age: 42 years Established ROI-based cortical thickness and subcortical volume SuStaIn 3 Subtype 1 (widespread): subcortical volume reduction
Subtype 2 (widespread): widespread reduction in cortical thickness and increase in GP volume
Subtype 3 (regional): widespread reduction in cortical thickness
No subtype differences in PANSS scores or proportion of treatment-resistant cases or proportion of patients prescribed clozapine.

Tensor-Based Morphometry Measures

Shi et al., 2023 (22) 673 (355 patients and 318 HCs); ∼52% male; ∼avg. age: 27 years Early stage ROI-based normative scores for TBM metric (Jacobian determinant) Sparse k-means clustering 2 Both subtypes showed widespread TBM value reduction in cortical and subcortical areas; subtypes differed in the degree of reduction, with subtype 2 having more extensive subcortical reductions; cerebellar TBM values were higher in subtype 1 (widespread) and lower in subtype 2 (widespread). No subtype differences in age, sex, illness duration, or antipsychotic dose; subtype 2 had a higher PANSS negative symptom and total score.

DTI Measures
Sun et al., 2015 (14) 223 (113 patients and 110 HCs); ∼51% male; ∼avg age: 23 years FEP ROI-based FA and MD Agglomerative hierarchical clustering 2 Subtype 1 (widespread): widespread reduction in FA coupled with increased MD in nearly all tracts
Subtype 2 (regional): mostly localized FA reduction in the SLF and MD increase in the CST
No subtype differences in age, sex, illness duration, or PANSS positive subscale scores; subtype 1 had higher PANSS negative subscale scores.

Detailed information and qualitative assessment of each study are provided in Table S1.

avg., average; CHIMERA, clustering of heterogeneous disease effects via distribution matching of imaging patterns; CSF, cerebrospinal fluid; CST, corticospinal tract; CV, cross-validation; DTI, diffusion tensor imaging; FA, fractional anisotropy; FEP, first-episode psychosis; GAF, Global Assessment of Function; GMC, gray matter concentration; GMV, gray matter volume; GP, globus pallidus; HC, healthy control; HYDRA, heterogeneity through discriminative analysis; ICA, independent component analysis; IDSCN, individual differential structural covariance network; MD, mean diffusivity; MSN, morphometric similarity network; MUSE, Multi-atlas Region Segmentation Utilizing Ensembles; opNMF, orthogonal projective non-negative matrix factorization; PANSS, Positive and Negative Syndrome Scale; PCA, principal component analysis; ROI, region of interest; SLF, superior longitudinal fasciculus; SuStaIn, subtype and stage inference; TBM, tensor-based morphometry; VBM, voxel-based morphometry; WMV, white matter volume.

a

The sample size reported is that of the primary/discovery sample; numbers reflect the sample included in the analyses, which may be smaller than the original study sample following quality control or other exclusions.

b

The clinical stage of patients is not operationally defined across studies; accordingly, the term first-episode psychosis is used when the sample is described as such in the original study, the term early stage is used for patients with a duration of illness <5 years as inferred from the sample demographics, and the term established stage is used for samples with longer illness duration.

c

Reductions or increases are typically referenced to the healthy control group.

d

The associated features reported are those examined in each original study.

Table 2.

Structured Qualitative Appraisal of Methodological Domains and Validation Practices Across Studies

Study Clinical Sample Representativeness and Provenance Transparency and Appropriateness of Preprocessing Pipelines Handling of Confounds Clustering Model Choice Evaluation of Cluster Stability External Validation, Generalizability Availability of Code and Parameter Settings
Gupta et al., 2017 (24) Multicohort sample from shared, deidentified legacy datasets T1-weighted MPRAGE sequence, on 1.5T and 3T scanners (site-specific parameters reported); preprocessing with affine normalization to MNI, reslicing to 2 × 2 × 2 mm, segmentation in SPM5, and 10-mm FWHM smoothing prior to VBM Age, sex, and site regressed voxelwise from the images B-ICA; no multialgorithm analyses reported No internal validation reported No external validation reported No publicly available code or scripts
Dwyer et al., 2018 (16) Multicohort sample from shared, deidentified legacy datasets Discovery imaging: T1-weighted 5-echo MPRAGE on 3T Siemens TIM; site-specific parameters reported; preprocessing with VBM with nonlocal-means filtering, tissue segmentation, Markov random field modeling; DARTEL normalization to MNI (3 mm isotropic) and 3 mm smoothing; images scaled to global GMV Age and sex were corrected using a regression model trained on healthy control participants. Fuzzy c-means with consensus clustering; criterion for optimal cluster number not reported; no multialgorithm analyses reported Subtype stability was assessed using repeated nested cross-validation with linear SVMs, permutation-based p values, and McNemar tests for group comparisons External validation in an independent sample using distance from the discovery cluster centroids No publicly available code or scripts
Honnorat et al., 2019 (30) Single-site cohort; not publicly available T1-weighted 3D MPRAGE sequence on a 1.5T scanner (parameters reported); preprocessing in SPM followed by MUSE segmentation to extract 80 ROI volumes Age, sex, and height were incorporated into the CHIMERA function to match patients and control participants on these covariates. CHIMERA; no cluster-number selection indices reported; no multi-algorithm analyses reported No internal validation metrics reported; sensitivity analyses for sex were included No external validation reported No publicly available code or scripts
Ma et al., 2019 (21) Single-site cohort; data not publicly available T1-weighted volumetric 3D SPGR sequence on a 3T scanner (detailed sequence parameters not reported); VBM preprocessing in SPM8 using DARTEL to derive GMV maps at 1.5 mm3 resolution, with regions defined using the AAL atlas (90 ROIs) Age, sex, and education are included as covariates in the gray matter volumes linear regression model K-means clustering using GMV of the control participants with distance = 1 − correlation; the optimal number of clusters chosen by mean IGP; no multialgorithm analyses reported 1000 replicate k-means runs per test with random initial centroids No external validation reported No publicly available code or scripts specific to this study
Chand et al., 2020 (19) PhenoM consortium subsample from 3 sites; data accessible on request T1-weighted images with diverse protocols (MPRAGE and BRAVO), scanners, and field strengths (1.5T and 3T); preprocessing was followed by MUSE segmentation into GMV, WMV, and CSF regions of interest Site effects in ROI volumes were estimated and removed using a linear model trained in healthy control participants; subsequently ROI volumes were adjusted for age and sex HYDRA; optimal cluster solution based on ARI; no multialgorithm analyses reported Subtype stability evaluated using permutation testing, split-sample methods, and leave-one-site-out validation No external validation reported HYDRA code is publicly available
Liu et al., 2021 (17) Multisite cohort comprising drug-naïve patients with FEP, chronic patients, and CHR individuals; data can be accessed upon request T1-weighted images acquired across sites on 3T and 1.5T scanners (site-specific parameters reported); VBM preprocessing in CAT12 (segmentation, DARTEL template generation, MNI normalization, 8 mm smoothing); regional GMV defined using the AAL2 atlas (94 ROIs); Network Template Perturbation used to estimate individualized structural covariance network (edgewise z scores relative to a control-derived reference network) Age, sex, education, and TIV were modeled as covariates in structural covariance estimation K-means with distance = 1 − correlation clustering applied to the top 20 Bonferroni-corrected deviant edges; cluster selection based on silhouette analysis; no multialgorithm analyses reported K-means repeated 100 times to derive the final clustering index Replication of the discovery subtype solution across additional independent FEP, chronic, and CHR datasets using the same 20-edge feature set No publicly available code or scripts specific to this study
Chew et al., 2022 (13) Single-site cohort; no data availability statement T1-weighted 3D MPRAGE sequence on a 3T scanner (parameters reported); preprocessing using standard FreeSurfer pipelines to derive regional morphometric measures Age, sex, and antipsychotic dose (CPZ equivalents) used as covariates in the clustering model HYDRA; optimal cluster solution based on ARI; no multialgorithm analyses reported; applied models pretrained in the Chand et al. study (19) No internal validation metrics reported because clustering was performed via application of pretrained models Explicitly designed as a replication of the study by Chand et al. (19) HYDRA code is publicly available
Dwyer et al., 2023 (18) PhenoM consortium subsample from 3 sites; none of these sites were part of the sample used by Chand et al. (19); data accessible on request T1-weighted images with diverse protocols (MPRAGE and BRAVO), scanners, and field strengths (1.5T and 3T); preprocessing was followed by MUSE segmentation into GMV, WMV, and CSF regions of interest; same pipeline as in Chand et al. (19) Age, sex, and antipsychotic dose (CPZ equivalents) used as covariates in the clustering model HYDRA; optimal cluster solution based on ARI; unsupervised k-means analysis used as a supplementary comparison No internal stability metrics for subtype discovery were reported because clustering was performed via application of pretrained models. Explicitly designed as a replication of the study by Chand et al. (19) HYDRA code is publicly available; additional code for subgroup assignment and signature computation is available via a dedicated GitHub repository.
Jiang et al., 2023 (20) Multisite cross-sectional (n = 9) and longitudinal (n = 2) cohorts; 5 datasets are public, and all others are accessible upon request T1-weighted images acquired using diverse protocols and scanners (site-specific parameters reported); preprocessing involved standard VBM via CAT12/SPM12; parcellation (AAL atlas) to extract 17 composite ROIs Age, sex, age2, and TIV were regressed out; site was modeled as a covariate; medication (CPZ equivalents) and illness stage were not used in the clustering model SuStaIn; Hopkins statistics used to evaluate clustering tendency; no multialgorithm analyses reported Subtype stability assessed using 2-fold cross-validation repeated 10 times; Dice coefficients used for subtype label consistency and Kendall’s tau for event sequence consistency, alongside silhouette values; additional stability checks included leave-one-site-out resampling and sensitivity analyses across atlases Analyses in independent longitudinal follow-up datasets Custom analysis code and the SuStaIn implementation are available at a dedicated GitHub.
Jiang et al., 2024 (11) Multicohort sample comprising 41 datasets; some datasets are public, and others are accessible upon request. T1-weighted images were acquired using diverse protocols and scanners (site-specific parameters reported); preprocessing involved standard VBM via CAT12/SPM12; parcellation (AAL atlas) to extract 17 composite ROIs Age, sex, age2, and TIV were regressed out; site harmonization was implemented with ComBat. SuStaIn; Hopkins statistics used to evaluate clustering tendency; no multialgorithm analyses reported Subtype stability was assessed via non-overlapping 2-fold cross-validation repeated 10 times; the Dice coefficient was used to quantify subtype label consistency Independent samples stratified by ancestry; subtype agreement with the original model was quantified as the proportion of unseen individuals who retained the same subtype label Custom analysis code and the SuStaIn implementation are available at a dedicated GitHub.
Chai et al., 2023 (12) Multicohort site; data publicly available T1-weighted on different 3T scanners with site-specific sequence parameters reported; preprocessing used a VBM pipeline in SPM12 to produce GMV maps smoothed at 8-mm FWHM; PCA applied to extract the GMV input features Age, sex, and TIV were regressed out; medication status was not modeled during clustering; ComBat was applied for site harmonization K-means and hierarchical clustering used to derive subtypes from GMV; model selection used ARI and the Calinski–Harabasz index Subtype stability was assessed using randomization with 5-fold cross-validation; repeatability assessed with the Jaccard similarity coefficients and maximum subtype-label probability across folds; cross-site evaluation across the datasets No external validation reported Critical scripts for clustering analyses are publicly available on a dedicated GitHub.
Wen et al., 2022 (29) Multisite samples from the Phenom, UK Biobank, and ADNI cohorts; data are available by request. T1-weighted on different 3T scanners with site-specific sequence parameters reported; GMV maps were reduced using orthogonal projective NMF to produce subject-level loading coefficients across components; the clustering algorithms were applied to the loading coefficients Age and sex were controlled via linear regression, with beta coefficients estimated solely on the healthy control participants; site harmonization through ComBat-GAM MAGIC and HYDRA used to derive subtypes; optimal cluster solution determined with ARI Subtype stability was assessed using cross-validation with 100 repeated stratified random splits No external validation in an independent sample; no quantitative metric of agreement between MAGIC- and HYDRA-derived subtypes MAGIC code is publicly available via a dedicated GitHub.
Yakimov et al., 2024 (23) Single-site cohort; publicly available T1-weighted acquired at 3T (specific parameters reported); images processed through standard FreeSurfer pipelines to yield measures of the choroid plexus and lateral ventricles; left and right hemisphere volumes were summed Age and sex were adjusted by residualizing choroid plexus and lateral ventricle volumes via linear regression. K-means with distance metric; the optimal solution was determined using the elbow method; no multialgorithm analyses reported No internal validation reported No external validation reported The software used is publicly available.
Yao et al., 2023 (28) Single-site cohort; data available upon request T1-weighted 3D FSPGR sequence on a 3T scanner (specific parameters reported); preprocessing in FreeSurfer; 308-region parcellation to derived measures of surface area, cortical thickness, gray matter volume, Gaussian curvature, mean curvature used to construct a 308 × 308 MSN; regional MSN strength used as input features Age, sex, and TIV were linearly regressed from MSN strength HYDRA; optimal cluster solution determined with ARI; no multialgorithm analyses reported Subtype stability was assessed using 10-fold cross-validation with 50 iterations. Independent replication in an external dataset; the similarity between the clusters in the original and replication samples was evaluated by correlating the spatial pattern of regional case-control t statistics HYDRA code is publicly available; no other study-specific scripts are openly available.
Xiao et al., 2022 (15) Single-site cohort; data available upon request T1-weighted spoiled gradient recall sequences on 3T scanners (site-specific parameters reported); preprocessing using standard FreeSurfer pipelines to yield measures of regional cortical area, thickness, and subcortical volume; PCA was applied to the z scores of these features followed by computation of pairwise Mahalanobis distances in the PCA space to generate a dissimilarity matrix for clustering Age, sex, and education were regressed from the neuroanatomical feature matrix. PDC; no formal optimization index reported; no multialgorithm analyses reported No internal validation metrics reported External validation using an independent publicly available cohort from the B-SNIP study; the similarity of the subtypes’ neuroanatomical alteration patterns across cohorts quantified with Pearson correlation The PDC algorithm code and other study-specific scripts are not provided.
Sone et al., 2024 (32) Multisite data from 3 cohorts T1-weighted BRAVO and FSPGR sequences on 3T scanners (site-specific parameters reported); preprocessing used standard FreeSurfer pipelines to yield measures of regional cortical area and thickness; 28 ROI z scores (selected from previous mega-analyses) were used for clustering Subcortical volumes corrected for TIV and cortical thickness and subcortical volumes corrected for age and sex; no site harmonization SuStaIn; no multialgorithm analyses reported Subtype stability assessed via 10-fold cross-validation and reproducibility analysis repeating the model in each site cohort separately No independent external validation The SuStaIn implementation is available at a dedicated GitHub.
Shi et al., 2023 (22) Multisite dataset collected from 7 cohorts; data available on request T1-weighted images were acquired using similar sequences in different 3T scanners (site-specific parameters reported); preprocessing used CAT12 pipelines with registration to the MNI152-2009c template; TBM features were derived as Jacobian determinants from the nonlinear deformation fields (local expansion/contraction relative to the template); individual deviation (W score) maps were computed for patients from voxelwise normative models; dimensionality reduction used the Brainnetome Atlas to create 273-dimensional regional W score vectors per patient, which were used for clustering. Normative modeling used a voxelwise general linear model fit in control participants with age, sex, and site as covariates; W scores for patients were derived from model residuals scaled by the control residual dispersion Sparse k-means tuning: the sparsity hyperparameter(s) was chosen using a permutation-based gap statistic (100 permutations per candidate s); cluster number selection was determined using the Calinski-Harabasz index and the silhouette coefficient; no multialgorithm analyses reported Subtype stability was assessed by computing Pearson correlations between the sparse k-means feature-weight vectors obtained under different partitioning strategies (merged dataset and leave-one-site-out runs) External validation used an independent sample; cross-sample generalization was based on the accuracy of an RBF-SVM trained on the discovery cohort and tested on the validation sample. Standard toolboxes and specific clustering parameters are detailed.
Sun et al., 2015 (14) Single-site cohort DTI sequence with 20 diffusion-encoding directions on 3T scanner (scanning parameters reported); preprocessing using standard pipelines in FSL followed by Automated Fiber Quantification for tractography; FA and MD values were derived from 36 tracks per participant and used in clustering – Used agglomerative hierarchical clustering; optimal cluster number was determined using silhouette, Dunn, and connectivity indices Stability was assessed using a subsampling No external validation reported Publicly available software packages; no repository for the study-specific scripts

Additional information on all studies is shown in Table 1.

AAL, Automated Anatomical Labeling; ADNI, Alzheimer’s Disease Neuroimaging Initiative; ARI, adjusted Rand index; B-ICA, biclustered independent component analysis; BRAVO, Brain Volume; CAT12, Computational Anatomy Toolbox12; CHIMERA, clustering of heterogeneous disease effects via distribution matching of imaging patterns; CPZ, chlorpromazine; CSF, cerebrospinal fluid; DARTEL, diffeomorphic anatomical registration using exponentiated lie algebra; DTI, diffusion tensor imaging; FA, fractional anisotropy; FEP, first episode psychosis; FSL, FMRIB Software Library; FSPGR, fast spoiled gradient-echo sequence; FWHM, full width at half maximum; GAM, generalized additive model; GMV, gray matter volume; HYDRA, heterogeneity through discriminative analysis; IGP, in-group proportion; MAGIC, multi-scale semi-supervised clustering; MD, mean diffusivity; MNI, Montreal Neurological Institute; MPRAGE, magnetization-prepared rapid acquisition with gradient-echo; MSN, morphometric similarity network; MUSE, Multi-atlas Region Segmentation Utilizing Ensembles; PCA, principal component analysis; PDC, peak density clustering; RBF-SVM, radial basis function-support vector machine; ROI, region of interest; SPGR, spoiled gradient-recalled echo; SuStaIn, subtype and stage inference; TBM, tensor-based morphometry; TIV, total intracranial volume; VBM, voxel-based morphometry; WMV, white matter volume.

Samples

The studies identified examined a broad range of populations with regard to sample size, demographic and clinical characteristics, and geographic origin (Table 1). The size of the patient samples across studies was small to modest (range 71–1124). The study by Jiang et al. (11) is a notable outlier because it included more than 4000 patients from 41 independent sites. Some studies focused specifically on first-episode or early-stage patients (12, 13, 14, 15), while the majority included established cases or mixed samples. Generally, the samples involved young to early middle-aged adults, with the average ages across studies clustered between 25 and 35 years. The percentage of male participants across samples varied from approximately 49% (15) to 74% (16), with most studies reporting a male predominance (the modal range across studies was 60%–70% male) (Table 1).

The clinical stage of patients was explicitly defined in the studies that focused on first-episode psychosis (FEP) (14,15,17,18). The remaining studies included cases that could be characterized either as early stage (defined here as having average illness duration of <5 years) or established (defined here as having longer illness duration) or a mixture of both. Most studies included patients with varied antipsychotic exposure, while few focused on drug-naïve patients (12,14,15). Several investigations were conducted in China (12, 13, 14, 15,17), while other studies relied on multisite datasets primarily from institutions based in North America and Europe (11,16,18, 19, 20).

Clustering Methods

Studies used a wide range of clustering methods, reflecting evolving computational strategies, but unsupervised approaches predominated (Figure 1A; Supplement). Among these, k-means and its variants were most common. Several studies used standard k-means (12,17,21, 22, 23), a centroid-based method that partitions individuals into k clusters by minimizing within-cluster variance. Because k-means assumes spherical clusters in Euclidean space, it was often preceded by dimensionality reduction, usually principal component analysis, to reduce noise and improve cluster separation. Sparse k-means was used in one study to select the most informative features (22), while fuzzy c-means assigned individuals partial membership across clusters to capture uncertainty in neuroanatomical data (16).

Figure 1.

Figure 1

Frequency of clustering algorithms and clinical associations of neuroanatomical subtype categories across studies. (A) Frequency of clustering algorithm usage across included studies (n = 18). Bars show the number of studies that applied each clustering approach to derive patient-level subtypes from neuroimaging measures. Bar segments indicate the imaging modality used as input to clustering: structural MRI (dark blue) and diffusion MRI (light blue). Where both modalities appear within a single bar, this indicates that the algorithm was used in studies drawing on different imaging modalities within the included literature. (B) Subtype categories and reported clinical associations across the included literature. Stacked bars show the total number of subtypes assigned to each neuroanatomical category (widespread, regional, preserved) based on the extent of structural abnormalities described in the primary studies. Within each category, segments indicate whether any clinical association with subtype membership was reported (dark green) or whether no clinical association was reported (light green). Numbers within the bars denote the count of subtypes in each segment. AHC, agglomerative hierarchical clustering; B-ICA, biclustered independent component analysis; CHIMERA, clustering of heterogeneous disease effects via distribution matching of imaging patterns; HYDRA, heterogeneity through discriminative analysis; MAGIC, multi-scale semi-supervised clustering; MRI, magnetic resonance imaging; PDC, peak density clustering; SuStaIn, subtype and stage inference.

Other studies used agglomerative hierarchical clustering (12,14), which iteratively merges similar data points to form a tree-like structure from which subtypes are defined at a chosen level. Less common unsupervised methods included biclustered independent component analysis (24), which simultaneously clusters subjects and brain regions but leaves many patients unassigned, and peak density clustering (15,25), which identifies cluster centers as high-density points distant from other density peaks and allows flexible, nonspherical boundaries.

The most recent unsupervised method was subtype and stage inference (SuStaIn) (26), used in the 2 largest studies (11,20). SuStaIn identifies subtypes and infers disease progression from cross-sectional data using event-based models. It defines subtype-specific sequences of regional brain changes and probabilistically assigns patients to positions along these trajectories based on deviation scores compared with healthy control participants. However, SuStaIn can accommodate only a limited number of variables, typically fewer than 20, because higher-dimensional input greatly increases computational complexity.

Among semisupervised methods, heterogeneity through discriminative analysis (HYDRA) (27) was most common (13,18,19,28). HYDRA uses multiple support vector machine classifiers and control data to optimize patient-control separation and identify disease-related heterogeneity. Multi-scale semi-supervised clustering (MAGIC) (29), a related approach, differs by first reducing voxelwise data into sparse, interpretable components using orthogonal projective non-negative matrix factorization before clustering. Clustering of heterogeneous disease effects via distribution matching of imaging patterns (CHIMERA) (30,31) is a registration-based framework that models heterogeneity by learning nonlinear transformations from healthy to disorder-related imaging patterns and assigns patients to the cluster whose transformation best reconstructs their neuroimaging profile.

Neuroimaging Input Features

The studies included predominantly relied on structural magnetic resonance imaging (sMRI) data, not because the review was restricted to this modality but because the studies identified through our search did not use other neuroimaging approaches; only one study incorporated diffusion tensor imaging (DTI) measures (14) (Figure 1A).

Across studies, 3 main types of features were used. Observed measures included regional gray matter volume or concentration derived from parcellations based on atlases such as MUlti-atlas region Segmentation utilizing Ensembles (MUSE), Automated Anatomical Labeling (AAL), Brainnetome, or FreeSurfer segmentations (Table 1). Deviation measures, calculated in several studies (11,15,20,22,32), compared patient-level neuroimaging values to healthy reference distributions to yield individualized deviation scores, which were then entered into clustering algorithms. Component-based measures, used less frequently, involved higher-order anatomical features such as source-based morphometry components (24), structural covariance network features (17), morphometric similarity networks (28), and orthogonal projective non-negative matrix factorization components (29). The sole DTI study (14) extracted fractional anisotropy (FA) and mean diffusivity (MD) from 18 major white matter tracts.

Subtype Number and Neuroimaging Patterns

The subtypes differed in the magnitude, direction, and spatial distribution of anatomical alterations but could be grouped into 3 descriptive categories (Table 1, Figure 2). Widespread subtypes showed broadly distributed deviations across multiple cortical lobes and/or several subcortical structures, typically dominated by reductions in morphometric measures relative to controls or normative expectations. Regional subtypes showed deviations that were confined to a limited set of anatomically circumscribed regions (for example, restricted to 1 lobe, a small number of cortical parcels, or 1–2 subcortical structures). Preserved subtypes showed minimal or no detectable deviations on the imaging features used for clustering. These labels are descriptive and do not imply shared neurobiological mechanisms.

Figure 2.

Figure 2

Clustering algorithms and resulting subtype categories across included studies. The Sankey diagram links each clustering algorithm (left) to the subtype category assigned in this review (right): widespread, regional, or preserved based on the extent of structural abnormalities described in the primary studies. The figure illustrates that the same algorithm can yield different subtype patterns across cohorts and feature sets and that similar subtype categories can be obtained with different algorithms. Each flow represents the number of subtypes reported using a given algorithm that fell into each category, and flow width is proportional to this count. Colors correspond to the algorithm on the left. Totals on the right (n = 23 widespread; n = 14 regional; n = 10 preserved) indicate the overall number of subtypes classified into each category across all studies and algorithms (i.e., counts of subtype solutions). AHC, agglomerative hierarchical clustering; B-ICA, biclustered independent component analysis; CHIMERA, clustering of heterogeneous disease effects via distribution matching of imaging patterns; HYDRA, heterogeneity through discriminative analysis; MAGIC, multi-scale semi-supervised clustering; PDC, peak density clustering; SuStaIn, subtype and stage inference.

Widespread subtypes were the most frequently reported. These subtypes showed widespread reduction in cortical and subcortical gray matter volume or cortical thickness in patients compared with normative expectations (11, 12, 13,15,16,18, 19, 20,29,32) or in morphometric similarity indices (28). Across studies, abnormalities typically encompassed regions within the frontal cortex, superior and middle temporal gyri, thalamus, insula, and hippocampus, although the magnitude and the specific localization of the regional effects varied markedly between studies. Widespread subtypes were observed in both patients with established illness (12,13,15,16,19,29,32) and first-episode/early-stage cases (15,18,28). In the only DTI study (14), the widespread subtype demonstrated white matter abnormalities in FA and MD across major association, commissural, and projection tracts.

Regional subtypes were reported in a smaller subset of studies and were characterized by anatomically restricted deviations in cortical and subcortical regions; the specific regions implicated and the nature of the alterations showed little consistency across studies or across the clinical stage of the samples involved (11, 12, 13, 14, 15,20,21,24,30,32). Examples include subtypes with reduced gray matter concentration in the insula, superior temporal gyrus, and inferior frontal gyrus or in the superior, middle, and medial frontal gyri (24); reduced gray matter volume in the thalamus, anterior cingulate, and superior temporal regions (30); or even more regionally restricted gray matter changes in the supramarginal gyrus, precuneus, and angular gyrus (21). Other reported patterns included subtypes with selective hippocampal volume loss (15) or volume reductions in other specific subcortical regions (11,20,32) or focal basal ganglia volume enlargement (13) or localized prefrontal and premotor volume increases (12). The only DTI study (14) identified a regional subtype with localized FA reduction in the superior longitudinal fasciculus coupled with increased MD in the corticospinal tract. One study also described a subtype defined by choroid plexus and lateral ventricle enlargement, although its broader anatomical profile was unclear due to the limited structural features assessed (23).

A small number of studies reported subtypes with largely preserved neuroanatomical profiles (15,17, 18, 19). These included patients with established illness (19) or FEP (18) showing preserved gray and white matter volume. Other studies identified preserved subtypes based on preserved covariance patterns in gray matter volume (15) or structural dissimilarity measures (17), with such profiles observed in both early-stage and established cases.

The SuStaIn-based subtypes identified by Jiang et al. (11,20) require further clarification. These studies yielded the 2 subtypes that seemed to differ primarily in the sequence and spatial emphasis of progressive gray matter loss. Both subtypes began as regionally circumscribed; in one, early changes were concentrated in prefrontal regions before extending to widespread cortical and subcortical areas, and in the other, subcortical alterations predominated initially and subsequently progressed to a more widespread pattern. Using the same method in a different sample, Sone et al. (32) identified 3 subtypes defined by initially selective subcortical reductions, widespread cortical thinning, or a combined profile that also included pallidal volume increase.

As illustrated in Figure 2, each algorithm used in more than one study produced subtypes spanning multiple neuroanatomical categories, indicating that subtype inconsistency persists even when the same clustering framework is applied across different cohorts.

Associated Demographic and Clinical Features of Neuroanatomical Subtypes

Subtype differences in demographic and clinical characteristics and functional outcomes were examined. However, there was substantial interstudy variability in the selection and reporting of these associated features as summarized in Table 1.

Age and sex were evaluated in nearly all studies. The vast majority found no significant differences between subtypes, regardless of whether cohorts comprised early-stage (15,17,18,22,28) or established (11,13,14,19,23,29) patients. One study reported that the widespread subtype was older, and the preserved subtype comprised more females (30), but no clinical differences accompanied these demographic observations.

The relationship between neuroanatomical subtype membership and clinical characteristics is summarized in Table 1 and Figure 1B. Of the 18 studies, 6 reported no significant differences between subtypes on any clinical measure examined (11,15,17,19,28,32). Across all studies, Positive and Negative Syndrome Scale (PANSS) symptom scores showed no subtype differences in 10 of 18 studies (11,13,15,17,19,23,28, 29, 30,32), differences in illness duration were similarly absent in 10 studies (11,15,17,19,22,23,28, 29, 30,32), and antipsychotic dose showed no subtype differences in 8 studies (11,15,16,19,22,23,29,30). Among the 11 studies that reported at least one clinical association, the nature and magnitude of differences varied considerably. When symptom differences were detected, they reflected a severity gradient aligned with degree of structural abnormality; widespread subtypes showed higher PANSS symptom scores (16,18,20, 21, 22,24), longer illness duration (13,16), and lower global functioning (29), while preserved subtypes generally showed more favorable clinical characteristics (11,18,20). Other reported associations included differences in age of illness onset (12,16), antipsychotic type or dose (12,18), and greater cognitive impairment in the subtype with the most pronounced structural changes (23). However, the predominant finding is that neuroanatomically distinct subgroups are largely indistinguishable on standard clinical measures.

Appraisal of Study Quality and Validation Practices

In the absence of a standard approach for rating the quality in clustering studies, we applied a structured framework focusing on key methodological domains. First, clinical sample representativeness is critical, as clusters reflect the population sampled; highly selected or narrowly defined samples may produce findings that do not generalize. Second, transparency and appropriateness of preprocessing pipelines are essential because clustering of high-dimensional neuroimaging data is highly sensitive to these steps. Third, handling of potential confounds (such as age, site, or treatment exposure) is also important, as uncorrected confounds can drive results independent of disease-related neurobiology. Fourth, clustering model choice and specification is a major concern because subtypes may reflect algorithmic preferences more than the underlying data structure. Fifth, cluster stability indicates whether similar clusters recur across repetitions or subsamples; unstable solutions suggest overfitting to noise or idiosyncrasies of a particular split. Sixth, external validation and generalizability reduce the risk of dataset-specific artifacts. Seventh, availability of code and parameters supports transparency and replication. Using this framework, we systematically evaluated the methodological rigor of each study and summarized these assessments in Table 2.

Validation procedures varied markedly across studies (Table 2). Internal stability was assessed in approximately half of the included studies, most often by repeating clustering across random initializations or subsamples and quantifying agreement with metrics such as the Jaccard similarity coefficient (12), permutation-based adjusted Rand index (19,28,29), Pearson correlation of feature-weight vectors (22), or Dice coefficients and Kendall’s tau across repeated cross-validation and leave-one-site-out resampling (11,20). The remaining studies reported no formal stability analysis or relied on qualitative inspection of cluster separation. External validation was less common. Two studies (13,18) were designed as attempts to replicate a previously reported subtype solution (19), whereas a further six studies evaluated subtype structure in an independent cohort (15, 16, 17,22,28), including Jiang et al. (11), who examined generalizability across ancestry-stratified independent samples. Only one study (20) assessed the temporal stability of subtype assignments. Overall, internal stability was addressed more often than external or longitudinal validation.

Discussion

This review synthesized recent neuroimaging-based clustering studies aimed at defining biologically grounded subtypes of schizophrenia, focusing on approaches that assign subtypes at the individual level. Across studies, there was substantial variability in sample characteristics, neuroimaging features, and clustering methodologies. Although the specific brain regions implicated varied considerably between studies, 3 broad conceptual patterns emerged representing subtypes with either widespread or more regional abnormalities and subtypes with preserved neuroimaging profiles. Associations with demographic and clinical variables were uncommon, but when present, they tended to link the widespread subtypes to greater clinical severity.

Methodological Considerations

Several methodological factors warrant consideration when interpreting this subtype literature. The first is sample size. Most studies were modest in scale, with only one including approximately 4000 patients (11). Limited sample size reduces power to detect smaller subgroups or those defined by subtle anatomical differences. This is especially relevant in schizophrenia, where case-control neuroimaging effects are themselves typically modest (8,9,33), implying that variation between putative subtypes is likely to be even smaller.

A second issue concerns the imaging features used for clustering. Most studies relied on sMRI, reflecting its broad availability and the maturity of pipelines for estimating cortical thickness, surface area, and subcortical volume. However, these macroscale measures may not capture the aspects of neurobiology most relevant to schizophrenia. Moreover, the features extracted from sMRI images varied substantially across studies, reflecting differences in parcellation schemes, dimensionality reduction methods, and normalization procedures, each of which can materially affect the resulting subtype solution. In addition to conventional sMRI measures, several studies used higher-order sMRI representations, including source-based morphometry components (24), structural covariance dissimilarity matrices (17), morphometric similarity network indices (28), and orthogonal projective non-negative matrix factorization components (29). These metrics are often less anatomically transparent, and in this review, their selection was seldom justified and did not seem to yield subtypes that were more robust or clinically informative.

The third issue influencing interstudy consistency in subtype solutions is the heterogeneity of the clustering methods. Most studies used conventional unsupervised approaches, including k-means, hierarchical clustering, or sparse and fuzzy variants. While computationally efficient, these algorithms impose simplifying assumptions about data geometry (e.g., spherical clusters) and similarity (e.g., Euclidean distance) that may be poorly suited to complex high-dimensional neuroimaging data. As a result, subtype solutions may be overly influenced by noise or by dominant global features. Semisupervised methods such as HYDRA and MAGIC introduce a different set of assumptions such as the separability of patient subtypes from control participants using linear or kernel-based decision boundaries that may not hold in practice. These approaches remain sensitive to imbalances in confounding variables (e.g., age, sex, site, or scanner effects) unless carefully addressed. In addition, most methods used yield hard subtype assignments without estimates of membership certainty, limiting the ability to assess reliability of classifications or to identify patients with ambiguous or unstable subtype membership.

The 2 largest studies in this review (11,20) used SuStaIn, but the assumptions of this approach require caution. Because the number of possible progression sequences increases exponentially with each added feature, SuStaIn cannot accommodate high-dimensional input and must be restricted to a small set of preselected regions. In these studies, regions were chosen from univariate case-control differences, which may bias subtype solutions toward the strongest diagnosis-related effects. SuStaIn also relies on normative deviation scores and is therefore sensitive to control group composition; unrecognized recruitment biases may distort these estimates and affect subtype assignment. In addition, it infers longitudinal progression from cross-sectional data, assuming that between-patient variation reflects a fixed sequence of accumulating pathology. In non-neurodegenerative disorders, however, such variation may instead reflect stable trait-like deviations present before illness onset.

Within-algorithm comparison, illustrated in Figure 2, further demonstrates that inconsistency is not driven by algorithm choice alone. HYDRA was applied in 4 studies (13,18,19,28), with a closely related framework (MAGIC) used in a fifth (29). Despite sharing the same semisupervised discriminative approach, subtype number ranged from 2 to 4 across these studies, and neuroanatomical profiles were inconsistent; some identified a widespread/preserved dichotomy, whereas others yielded a widespread/regional pattern (Figure 2). The same applied to the 3 studies using SuStaIn (11,20,32); 2 studies produced broadly convergent 2-subtype solutions, whereas the third identified 3 subtypes with a partially distinct configuration. K-means and its variants, the most frequently used family of methods, produced the most heterogeneous results in both subtype number and neuroanatomical profiles across studies (Figure 2). The overall pattern suggests that interstudy inconsistency reflects the collective influence of differences in sample composition, input feature selection, confound handling, and clustering framework. This observation reinforces the case for standardizing these upstream analytic decisions, encompassing harmonized feature extraction, explicit confound control, and prespecified sensitivity analyses across algorithm families, before the independent contribution of algorithmic diversity to subtype inconsistency can be meaningfully evaluated.

Clinical Relevance and Interpretability of Neuroimaging-Based Subtypes

The central finding of this review is that clustering reliably detects neuroanatomical heterogeneity, but the resulting subtypes show limited reproducibility across studies and little meaningful differentiation on clinical measures. These 2 observations are unlikely to be independent.

If neuroimaging-based clustering were identifying biologically distinct disease entities, one would expect subtype profiles to be both anatomically consistent and clinically differentiated. Their joint absence provides the strongest indication that current subtype solutions do not map onto discrete underlying mechanisms. Even within broad categories such as widespread or regional alterations, the regions involved and the direction of deviations varied substantially. This anatomical variability limits confidence in the validity of the subtypes identified. Without well-defined and reproducible anatomical profiles, it is challenging to delineate the underlying biological mechanisms of each subtype and even more difficult to develop them into constructs with clinical applicability.

These findings underscore the difficulties in distinguishing true biological subtypes from variation along a shared continuum. If schizophrenia comprised biologically distinct subtypes, clustering would be expected to recover consistent, orthogonal neuroanatomical patterns linked to distinct clinical profiles. The absence of such evidence suggests that the heterogeneity described in schizophrenia in terms of symptoms, severity, or course may not map onto separable morphometric processes. Instead, clinical variability may reflect different expressions of shared underlying biology shaped by development, environmental exposures, or co-occurring conditions. It is also possible that schizophrenia is causally heterogeneous but that distinct pathogenic pathways converge on similar structural brain changes because macroscale morphometric features may lie relatively far downstream from primary disease mechanisms. Alternatively, limitations may be inherent to macroscale neuroimaging features whose continuous and highly collinear nature may favor dimensional rather than categorical solutions. More fine-grained structural measures, or modalities more closely linked to molecular processes, may prove better suited to resolving meaningful biological variation.

The broader clustering literature in schizophrenia spectrum and related disorders extends beyond neuroimaging-only studies. Data-driven clustering has been applied to symptoms, cognition, and multimodal combinations of imaging, clinical, and biological measures, often in transdiagnostic samples (5,34, 35, 36, 37, 38). In principle, multimodal and transdiagnostic approaches may integrate complementary dimensions of variation but introduce major practical and interpretive challenges. Domains differ in scale, reliability, and missingness, increasing the risk that clusters reflect measurement artifacts rather than biology. High-variance domains can dominate distance-based methods, while correlated features across domains may overrepresent the same signal. Interpretation is also more difficult because clusters may conflate illness stage, treatment exposure, or cognitive impairment. Transdiagnostic sampling adds further ambiguity, as clusters may track diagnostic composition or severity gradients rather than disease-relevant subtypes and often generalize poorly when case mix shifts across sites.

Conclusions and Future Directions

Current evidence does not yet establish whether macroscale neuroimaging features can define subtypes of schizophrenia that are either clinically or biologically meaningful. The anatomical inconsistency and limited clinical differentiation documented in this review define the evidential threshold that future subtyping studies must meet. In this context, the review serves to delineate the methodological and interpretive hurdles that must be addressed before subtype frameworks can be credibly considered for clinical or research use. A solution that is reproducible across cohorts but not clinically differentiated or clinically differentiated but anatomically inconsistent would not provide evidence for biologically grounded stratification.

A constructive way forward might be to define a set of evidential criteria, alongside methodological expectations, that can serve as benchmarking standards for studies aimed at subtyping schizophrenia and other clinical populations. These would include transparent and harmonized feature extraction and preprocessing, explicit identification and control of key confounds, reporting of cluster assignment uncertainty and stability under resampling and site/scanner variation, and evidence that a subtype solution generalizes to independent cohorts and remains coherent under longitudinal follow-up. Where clinical relevance is claimed, it also requires evaluation against outcomes that were not used to derive the subtypes so that clinical utility is not conflated with cluster construction. In parallel, greater justification of algorithm selection would strengthen inference. At a minimum, this involves aligning the chosen method’s assumptions with the structure of the data and the scientific question and demonstrating that the subtype solution is not an artifact of a particular algorithm or parameter setting through prespecified sensitivity analyses and, where feasible, comparisons across algorithm families applied to the same inputs.

Acknowledgments and Disclosures

The authors report no biomedical financial interests or potential conflicts of interest.

Footnotes

Supplementary material cited in this article is available online at https://doi.org/10.1016/j.bpsgos.2026.100745.

Supplementary Material

Supplemental Text and Figure S1
mmc1.pdf (132KB, pdf)

References

  • 1.Sepulveda A.R., Hamilton S.R., Allegra C.J., Grody W., Cushman-Vokoun A.M., Funkhouser W.K., et al. Molecular biomarkers for the evaluation of colorectal cancer: Guideline from the American Society for Clinical Pathology, College of American Pathologists, Association for Molecular Pathology, and the American Society of Clinical Oncology. J Clin Oncol. 2017;35:1453–1486. doi: 10.1200/JCO.2016.71.9807. [DOI] [PubMed] [Google Scholar]
  • 2.Cardoso F., Kyriakides S., Ohno S., Penault-Llorca F., Poortmans P., Rubio I.T., et al. Early breast cancer: ESMO Clinical Practice Guidelines for diagnosis, treatment and follow-up†. Ann Oncol. 2019;30:1194–1220. doi: 10.1093/annonc/mdz173. [DOI] [PubMed] [Google Scholar]
  • 3.Eshaghi A., Young A.L., Wijeratne P.A., Prados F., Arnold D.L., Narayanan S., et al. Identifying multiple sclerosis subtypes using unsupervised machine learning and MRI data. Nat Commun. 2021;12:2078. doi: 10.1038/s41467-021-22265-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Zhang X., Chou J., Liang J., Xiao C., Zhao Y., Sarva H., et al. Data-driven subtyping of Parkinson’s disease using longitudinal clinical records: A cohort study. Sci Rep. 2019;9:797. doi: 10.1038/s41598-018-37545-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Habtewold T.D., Rodijk L.H., Liemburg E.J., Sidorenkov G., Boezen H.M., Bruggeman R., Alizadeh B.Z. A systematic review and narrative synthesis of data-driven studies in schizophrenia symptoms and cognitive deficits. Transl Psychiatry. 2020;10:244. doi: 10.1038/s41398-020-00919-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Carpenter W.T., Jr., Kirkpatrick B. The heterogeneity of the long-term course of schizophrenia. Schizophr Bull. 1988;14:645–652. doi: 10.1093/schbul/14.4.645. [DOI] [PubMed] [Google Scholar]
  • 7.Howes O.D., Cummings C., Chapman G.E., Shatalina E. Neuroimaging in schizophrenia: An overview of findings and their implications for synaptic changes. Neuropsychopharmacology. 2023;48:151–167. doi: 10.1038/s41386-022-01426-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.van Erp T.G.M., Hibar D.P., Rasmussen J.M., Glahn D.C., Pearlson G.D., Andreassen O.A., et al. Subcortical brain volume abnormalities in 2028 individuals with schizophrenia and 2540 healthy controls via the ENIGMA consortium. Mol Psychiatry. 2016;21:547–553. doi: 10.1038/mp.2015.63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.van Erp T.G.M., Walton E., Hibar D.P., Schmaal L., Jiang W., Glahn D.C., et al. Cortical brain abnormalities in 4474 individuals with schizophrenia and 5098 control subjects via the Enhancing Neuro Imaging Genetics through Meta Analysis (ENIGMA) consortium. Biol Psychiatry. 2018;84:644–654. doi: 10.1016/j.biopsych.2018.04.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Kraguljac N.V., McDonald W.M., Widge A.S., Rodriguez C.I., Tohen M., Nemeroff C.B. Neuroimaging biomarkers in schizophrenia. Am J Psychiatry. 2021;178:509–521. doi: 10.1176/appi.ajp.2020.20030340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Jiang Y., Luo C., Wang J., Palaniyappan L., Chang X., Xiang S., et al. Neurostructural subgroup in 4291 individuals with schizophrenia identified using the subtype and stage inference algorithm. Nat Commun. 2024;15:5996. doi: 10.1038/s41467-024-50267-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Chai C., Ding H., Du X., Xie Y., Man W., Zhang Y., et al. Dissociation between neuroanatomical and symptomatic subtypes in schizophrenia. Eur Psychiatry. 2023;66 doi: 10.1192/j.eurpsy.2023.2446. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Chew Q.H., Prakash K.N.B., Koh L.Y., Chilla G., Yeow L.Y., Sim K. Neuroanatomical subtypes of schizophrenia and relationship with illness duration and deficit status. Schizophr Res. 2022;248:107–113. doi: 10.1016/j.schres.2022.08.004. [DOI] [PubMed] [Google Scholar]
  • 14.Sun H., Lui S., Yao L., Deng W., Xiao Y., Zhang W., et al. Two patterns of white matter abnormalities in medication-naive patients with first-episode schizophrenia revealed by diffusion tensor imaging and cluster analysis. JAMA Psychiatry. 2015;72:678–686. doi: 10.1001/jamapsychiatry.2015.0505. [DOI] [PubMed] [Google Scholar]
  • 15.Xiao Y., Liao W., Long Z., Tao B., Zhao Q., Luo C., et al. Subtyping schizophrenia patients based on patterns of structural brain alterations. Schizophr Bull. 2022;48:241–250. doi: 10.1093/schbul/sbab110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Dwyer D.B., Cabral C., Kambeitz-Ilankovic L., Sanfelici R., Kambeitz J., Calhoun V., et al. Brain subtyping enhances the neuroanatomical discrimination of schizophrenia. Schizophr Bull. 2018;44:1060–1069. doi: 10.1093/schbul/sby008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Liu Z., Palaniyappan L., Wu X., Zhang K., Du J., Zhao Q., et al. Resolving heterogeneity in schizophrenia through a novel systems approach to brain structure: Individualized structural covariance network analysis. Mol Psychiatry. 2021;26:7719–7731. doi: 10.1038/s41380-021-01229-4. [DOI] [PubMed] [Google Scholar]
  • 18.Dwyer D.B., Chand G.B., Pigoni A., Khuntia A., Wen J., Antoniades M., et al. Psychosis brain subtypes validated in first-episode cohorts and related to illness remission: Results from the PhenoM consortium. Mol Psychiatry. 2023;28:2008–2017. doi: 10.1038/s41380-023-02069-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Chand G.B., Dwyer D.B., Erus G., Sotiras A., Varol E., Srinivasan D., et al. Two distinct neuroanatomical subtypes of schizophrenia revealed using machine learning. Brain. 2020;143:1027–1038. doi: 10.1093/brain/awaa025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Jiang Y., Wang J., Zhou E., Palaniyappan L., Luo C., Ji G., et al. Neuroimaging biomarkers define neurophysiological subtypes with distinct trajectories in schizophrenia. Nat Mental Health. 2023;1:186–199. [Google Scholar]
  • 21.Ma L., Rolls E.T., Liu X., Liu Y., Jiao Z., Wang Y., et al. Multi-scale analysis of schizophrenia risk genes, brain structure, and clinical symptoms reveals integrative clues for subtyping schizophrenia patients. J Mol Cell Biol. 2019;11:678–687. doi: 10.1093/jmcb/mjy071. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Shi W., Fan L., Wang H., Liu B., Li W., Li J., et al. Two subtypes of schizophrenia identified by an individual-level atypical pattern of tensor-based morphometric measurement. Cereb Cortex. 2023;33:3683–3700. doi: 10.1093/cercor/bhac301. [DOI] [PubMed] [Google Scholar]
  • 23.Yakimov V., Moussiopoulou J., Roell L., Kallweit M.S., Boudriot E., Mortazavi M., et al. Investigation of choroid plexus variability in schizophrenia-spectrum disorders-Insights from a multimodal study. Schizophrenia (Heidelb) 2024;10:121. doi: 10.1038/s41537-024-00543-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Gupta C.N., Castro E., Rachkonda S., van Erp T.G.M., Potkin S., Ford J.M., et al. Biclustered independent component analysis for complex biomarker and subtype identification from structural magnetic resonance images in schizophrenia. Front Psychiatry. 2017;8:179. doi: 10.3389/fpsyt.2017.00179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Rodriguez A., Laio A. Machine learning. Clustering by fast search and find of density peaks. Science. 2014;344:1492–1496. doi: 10.1126/science.1242072. [DOI] [PubMed] [Google Scholar]
  • 26.Young A.L., Marinescu R.V., Oxtoby N.P., Bocchetta M., Yong K., Firth N.C., et al. Uncovering the heterogeneity and temporal complexity of neurodegenerative diseases with Subtype and Stage Inference. Nat Commun. 2018;9:4273. doi: 10.1038/s41467-018-05892-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Varol E., Sotiras A., Davatzikos C., Alzheimer’s Disease Neuroimaging Initiative Hydra: Revealing heterogeneity of imaging and genetic patterns through a multiple max-margin discriminative analysis framework. Neuroimage. 2017;145:346–364. doi: 10.1016/j.neuroimage.2016.02.041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Yao G., Zou T., Luo J., Hu S., Yang L., Li J., et al. Cortical structural changes of morphometric similarity network in early-onset schizophrenia correlate with specific transcriptional expression patterns. BMC Med. 2023;21:479. doi: 10.1186/s12916-023-03201-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Wen J., Varol E., Sotiras A., Yang Z., Chand G.B., Erus G., et al. Multi-scale semi-supervised clustering of brain images: Deriving disease subtypes. Med Image Anal. 2022;75 doi: 10.1016/j.media.2021.102304. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Honnorat N., Dong A., Meisenzahl-Lechner E., Koutsouleris N., Davatzikos C. Neuroanatomical heterogeneity of schizophrenia revealed by semi-supervised machine learning methods. Schizophr Res. 2019;214:43–50. doi: 10.1016/j.schres.2017.12.008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Dong A., Honnorat N., Gaonkar B., Davatzikos C. CHIMERA: Clustering of heterogeneous disease effects via distribution matching of imaging patterns. IEEE Trans Med Imaging. 2016;35:612–621. doi: 10.1109/TMI.2015.2487423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Sone D., Young A., Shinagawa S., Tsugawa S., Iwata Y., Tarumi R., et al. Disease progression patterns of brain morphology in schizophrenia: More progressed stages in treatment resistance. Schizophr Bull. 2024;50:393–402. doi: 10.1093/schbul/sbad164. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Voineskos A.N., Hawco C., Neufeld N.H., Turner J.A., Ameis S.H., Anticevic A., et al. Functional magnetic resonance imaging in schizophrenia: Current evidence, methodological advances, limitations and future directions. World Psychiatry. 2024;23:26–51. doi: 10.1002/wps.21159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Carruthers S.P., Van Rheenen T.E., Karantonis J.A., Rossell S.L. Characterising demographic, clinical and functional features of cognitive subgroups in schizophrenia spectrum disorders: A systematic review. Neuropsychol Rev. 2022;32:807–827. doi: 10.1007/s11065-021-09525-0. [DOI] [PubMed] [Google Scholar]
  • 35.Tornero-Costa R., Martinez-Millana A., Azzopardi-Muscat N., Lazeri L., Traver V., Novillo-Ortiz D. Methodological and quality flaws in the use of artificial intelligence in mental health research: Systematic review. JMIR Ment Health. 2023;10 doi: 10.2196/42045. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Clementz B.A., Sweeney J.A., Hamm J.P., Ivleva E.I., Ethridge L.E., Pearlson G.D., et al. Identification of distinct psychosis biotypes using brain-based biomarkers. Am J Psychiatry. 2016;173:373–384. doi: 10.1176/appi.ajp.2015.14091200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Dwyer D.B., Buciuman M.O., Ruef A., Kambeitz J., Sen Dong M., Stinson C., et al. Clinical, brain, and multilevel clustering in early psychosis and affective stages. JAMA Psychiatry. 2022;79:677–689. doi: 10.1001/jamapsychiatry.2022.1163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Lalousis P.A., Schmaal L., Wood S.J., Reniers R.L.E.P., Barnes N.M., Chisholm K., et al. Neurobiologically based stratification of recent-onset depression and psychosis: Identification of two distinct transdiagnostic phenotypes. Biol Psychiatry. 2022;92:552–562. doi: 10.1016/j.biopsych.2022.03.021. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Text and Figure S1
mmc1.pdf (132KB, pdf)

Articles from Biological Psychiatry Global Open Science are provided here courtesy of Elsevier

RESOURCES