Abstract
Psychiatric disorders are highly heterogeneous, shaped by intricate interactions among genetic, neurobiological, cognitive, social, and behavioral determinants. Traditional unimodal approaches often struggle to capture the complexity, limiting diagnostic precision and prognostic accuracy. The integration of diverse data modalities offers substantial promise for advancing precision psychiatry, enabling more nuanced, comprehensive, individualized insights into mental illnesses. Recent developments in artificial intelligence (AI), particularly machine learning (ML) and deep learning (DL), have accelerated the ability to synthesize and harness complex, multimodal datasets at scale. In this review, we systematically examine the integration and analysis of multimodal data sources through advanced AI models in psychiatry research, addressing 3 key dimensions: 1) primary data modalities, including texts (e.g., clinical assessment and questionnaires), neuroimaging, electrophysiological signals, audio and video recordings, and molecular and multi-omics profiles; 2) AI-based multimodal learning approaches, encompassing feature fusion strategies and state-of-the-art methodological paradigms ranging from conventional ML to DL and transformers; and 3) representative applications, spanning present-state characterization (diagnosis, subtyping, and stratification) to future-state prediction (risk, treatment response, and prognosis). Furthermore, we critique critical challenges that impede progress, including data-related barriers (unpaired modality, limited availability, and integration complexity) and model-related limitations (generalizability, interpretability, and clinical trustworthiness). Finally, we explore future opportunities, particularly multimodal large language models that offer unprecedented ingestion and reasoning capabilities across diverse modalities. We emphasize potential pathways (development of well-linked multimodal datasets, methodological innovation, and interdisciplinary collaboration) to realize the transformative potential of AI-empowered multimodal learning, thereby advancing personalized diagnostics, prognostics, and therapeutic strategies in precision psychiatry.
Psychiatric disorders constitute a leading cause of disease burden worldwide, contributing substantially to morbidity, disability, and premature mortality (1,2). In the United States, the prevalence of mental illness is more than 20% in the adult population and costs the economy over $280 billion annually, equivalent to 1.7% of gross domestic product (3,4). These conditions are inherently complex and heterogeneous, shaped by the intricate interplay of genetic, neurobiological, developmental, social, behavioral, and environmental factors (5–7). Traditional research approaches, often constrained to unimodal data, such as clinical assessments or neuroimaging, have provided important insights but remain insufficient to capture the multifaceted nature (8). As a result, clinical decision making in psychiatry continues to struggle with imprecision in diagnosis, prognosis, and treatment response, underscoring the urgent need for more comprehensive, data-driven approaches. The increasing availability of multimodal data sources, including electronic health records (EHRs), structural and functional neuroimaging, digital phenotyping from smartphones and wearables (9), audio and video recordings, as well as molecular and multi-omics profiles (genomics, transcriptomics, proteomics, etc.) (10), offers unprecedented opportunities to advance precision psychiatry. However, integrating and harnessing these heterogeneous, high-dimensional, and temporally dynamic data pose substantial analytical and computational challenges.
Recent advances in artificial intelligence (AI), including machine learning (ML), deep learning (DL), and large language models (LLMs), have demonstrated transformative potential in addressing these challenges. By enabling cross-modal feature fusion and latent pattern recognition, AI-based multimodal learning consistently outperforms unimodal approaches in accuracy, robustness, and generalizability across diverse psychiatric tasks (11). Notably, the advent of LLMs, especially multimodal variants such as OpenAI’s GPT series (version 4 and above) (12), further expands the capacity of ingesting and synthesizing multiple language and visual modalities. Adaptation strategies, such as fine-tuning, prompt engineering, and domain-specific pretraining on large-scale medical datasets, further enhance LLMs’ ability to represent psychiatric characteristics and improve performance in downstream tasks (13–16). Beyond methodological advances, multimodal learning has been extensively applied in psychiatry to elucidate biological underpinnings and predict clinical trajectories. These approaches enhance phenotypic resolution (17,18) and open new avenues for earlier and more precise detection (19,20), individualized prognostic stratification (21,22), and development of targeted therapeutic interventions (23,24) for mental disorders.
This review synthesizes the current landscape of AI-enabled multimodal data integration and prediction in the field of psychiatry. We summarize core data modalities and sources, examine key methodological frameworks, and highlight representative applications across psychiatric conditions. Furthermore, we discuss critical challenges, such as data availability and privacy and model interpretability and generalizability. We also discuss the opportunities brought by advanced AI techniques, particularly rapidly evolving multi-modal LLMs. By providing a comprehensive, in-depth, and forward-looking overview, we aim to provide insights to inform ongoing research efforts and facilitate the translational impact of AI-driven multimodal approaches in both psychiatric research and clinical practice.
METHODS
Literature Search
We conducted a comprehensive literature search across 5 major databases encompassing both biomedicine and computer science domains: PubMed, ACM Digital Library, Web of Science, EMBASE, and LENS (25). The search was restricted to studies published within the past 3 years (June 1, 2022–June 1, 2025). The search strategy included psychiatry- and AI-related medical subject heading terms and keywords (Supplement, section 1 for the complete query strategy) (Figure 1A).
Figure 1.

Large language model (LLM)–assisted literature selection and inclusion with human-in-the-loop review. (A) Overview of the pipeline from literature retrieval and deduplication to automated title/abstract screening, full-text human review, and final inclusion. (B) Multiagent screening workflow. Two locally deployed open-source LLMs (Llama 4 and Qwen3) independently assessed titles and abstracts using predefined criteria. Majority voting determined exclusion or provisional inclusion, while disagreements were resolved by GPT-4o as the arbiter. All provisionally included studies proceeded to human review.
Eligibility Criteria
Studies were included if they met all the following criteria: 1) explicitly addressed psychiatry-related disorders or psychiatric constructs (e.g., emotion, affect, or stress), 2) incorporated multimodal data sources, and 3) applied AI methodologies. Exclusion criteria included 1) non-original research (e.g., reviews, perspectives, commentaries, or case reports) and 2) studies using animals only.
Screening and Selection
To enhance screening efficiency, we developed a multiagent system comprising both open-source and proprietary LLMs with human-in-the-loop validation (Figure 1B). We adopted majority voting and arbitration mechanisms to achieve an optimal collective effect. Specifically, 2 locally deployed open-source LLMs, Llama 4 (26) and Qwen3 (27), serving as worker agents, independently evaluated study titles and abstracts against predefined inclusion and exclusion criteria. If both models excluded an article, it was discarded; if both models selected an article, it advanced to human review. In cases of disagreement, GPT-4o served as a leader agent (arbiter), with included studies proceeding to human assessment. This process yielded 240 candidate studies for full-text human review. Twelve co-authors carefully read a subset of the full articles (around 20/person), ensuring alignment with inclusion criteria. Final decisions were made through group discussion and consensus. Studies that were misclassified as multimodal or non–peer reviewed or otherwise failed to meet inclusion criteria were further excluded. Ultimately, 88 studies were retained for in-depth analysis and synthesis. The screening prompt for LLMs and inter-LLM and human-LLM agreement rates are detailed in Supplement, sections 2 and 3.
RESULTS
We synthesized the included studies across 3 main dimensions: data modalities, AI-based multimodal learning approaches, and psychiatry applications. Categorization within each dimension was performed by human-led review and summarization, followed by domain expert confirmation from both psychiatry and AI fields.
Data Modalities
The effectiveness of multimodal AI in psychiatry depends on access to high-quality, diverse data that reflect the multifaceted nature of psychiatric disorders. Our review of 90 studies identified primary data modalities that currently drive research in this area (Figures 2A and 3A; Table 1).
Figure 2.

Overview of multimodal data, artificial intelligence (AI) model architectures, and psychiatric applications. (A) Categories of multimodal data used in psychiatric research, including text, video, audio, imaging, time-series signals, and genetic/molecular data. (B) Representative AI architectures and fusion strategies, covering early, intermediate, and late fusion, along with machine learning, deep learning, transformer, and large language model–based models for multimodal learning. (C) Clinical applications spanning current-state characterization (diagnosis, subtyping, and phenotyping) and future-state prediction (risk, prognosis, and treatment response). CNN, convolutional neural network; GNN, graph neural network; KNN, k-nearest neighbor; RNN, recurrent neural network; SVM, support vector machine; ViT, vision transformer.
Figure 3.

Subcategory distributions across data modalities, artificial intelligence (AI) architectures, psychiatry applications, and psychiatric disorders. (A) Distribution of data modalities used across studies, encompassing audio, video, genetic and molecular (G&M) data, physiological and behavioral signals, imaging, and text. Subcategories highlight the diversity within each modality, such as speech and nonspeech audio, structural/functional neuroimaging, multi-omics data, and structured vs. unstructured clinical text. (B) Overview of AI architectures used, spanning traditional machine learning, Transformer-based models, multimodal hybrid models, and deep learning (DL) methods, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders (AEs), and graph-based networks. Subcategory counts reflect the heterogeneous modeling strategies adopted for multimodal fusion and prediction. (C) Psychiatric application domains, grouped into current-state characterization (diagnosis, subtyping, phenotyping) and future-state prediction (risk and onset prediction, prognosis, and treatment-response forecasting). (D) Distribution of psychiatric disorders studied, categorized based on medical subject heading concepts (132). Disorders include anxiety, neurodevelopmental conditions, mood disorders, schizophrenia spectrum disorders, and other less frequently studied conditions, reflecting variability in research focus across multimodal AI literature. Because some mood-related studies are not disorder focused (e.g., emotion identification), we classify them under mood within the category of mood disorders. LLM, large language model; ML, machine learning; ND, neurodevelopmental disorder; SSOPD, schizophrenia spectrum and other psychotic disorders.
Table 1.
Data Modalities in Psychiatry Research
| Modality | Sub-modality | Examples |
|---|---|---|
| Text | Structured table | diagnosis codes, medication, and lab results |
| Unstructured text | Clinical notes | |
| Semi-structured text | Clinical assessment, survey, questionnaire | |
| Image | Structural/Functional imaging | sMRI, fMRI, CT, X-ray, Ultrasound, PET |
| Video | Facial and body movement | Interview recordings, emotion recognition, motor assessment |
| Task performance | Psychomotor tasks, walking tests | |
| Audio | Speech audio | Interview audio, verbal recall, speech emotion analysis |
| Non-speech audio | Breathing, coughing, tone, or affective cues | |
| Signal | Motion and biosignal data | Behavioral/physiological time-series, such as eye tracking, motion tracking, accelerometry, and vital signs (e.g., heart rate) |
| Electrophysiological maps | EEG, MEG maps | |
| Genetic & Molecular Data | Genomic and genetic variants; Transcriptomics; Epigenomics; Proteomics; Metabolomics | Genomic and genetic variants (e.g., SNPs, CNVs, WGS/WES); Transcriptomics (e.g., RNA-seq, gene expression arrays); Epigenomics (e.g., DNA methylation, histone modifications); Proteomics (e.g., protein abundance via mass spectrometry); Metabolomics (e.g., metabolite concentrations from LC-MS, GC-MS) |
Textual data (55 studies) emerged as one of the most widely used modalities. Within the category, semistructured text, such as clinical assessment and questionnaires (23,28–32), was notably predominant (29 studies). These data were often integrated into tabular datasets for AI modeling or used directly as ground truth labels to reflect psychiatric status. Structured tabular data (12 studies), such as diagnosis codes (21,24,33), provided standardized descriptors that support interoperability across health systems. Unstructured text, such as clinical notes (34–36), offered context-rich information for patient treatment and clinical decision making. Together, these (semi)structured and unstructured text data form a valuable foundation for predictive modeling.
Imaging modalities (37 studies), particularly structural and functional neuroimaging (37–39), played an essential role in psychiatric AI research. Raw magnetic resonance imaging (MRI) data usually underwent preprocessing via automated pipelines such as Configurable Pipeline for the Analysis of Connectomes (40). Processed data were parcellated using anatomical atlases [e.g., Automated Anatomical Atlas (41)] to extract region of interest time series and construct functional connectivity matrices for downstream AI modeling.
Signal-based modalities (42 studies), such as accelerometry (18,42) and electrophysiological mapping (e.g., electroencephalography) (43–45), highlight their importance for capturing neural and behavioral dynamics. These signals underwent quality control to remove physiologically implausible values, followed by segmentation into non-overlapping windows for feature extraction. For example, accelerometry data segments were paired with the corresponding Hamilton Depression Rating Scale/Young Mania Rating Scale (46,47) clinical scores and incorporated as signal-modality input for downstream model training.
In addition, genetic and molecular data (19 studies) provided complementary microlevel insights (28,33,48). Genomic and genetic variants were the most frequently used, reflecting their growing role in diagnosis and prognosis. Other submodalities, such as RNA editing and metabolomics (49,50), were less frequently explored but represent promising avenues for future investigation. Audio (51–53) and video (22,54,55) modalities enable continuous, fine-grained behavioral monitoring that surpasses static imaging or textual reports. However, their adoption remains limited due to computational and resource-intensive processing requirements.
Regarding multimodal combinations, as shown in Figure 4, imaging-imaging pairings were the most prevalent (21 studies). Heterogeneous combinations, such as text-signal (15 studies), highlight a growing emphasis on integrating complementary information sources. Detailed breakdowns of data modality combinations are provided in Table S2. While psychiatric diagnosis continues to rely primarily on structured interviews and validated scales, often complemented by imaging and physiological signals to exclude organic etiologies, research increasingly explores multimodal combinations. Genetic and molecular data remain concentrated in specialized or developmental disorder contexts, whereas audio and video modalities show strong potential for emotion recognition and behavioral analysis.
Figure 4.

Data modality co-occurrence network in the selected studies. G&M, genetic and molecular.
AI-Based Multimodal Learning Approaches
In biological psychiatry, AI workflows typically follow a structured pipeline involving data acquisition, feature extraction, model training, and decision making. While unimodal approaches handle this process in a relatively straightforward manner, multimodal learning introduces additional complexity due to the need to integrate diverse sources at various stages of the pipeline (56,57).
Multimodal Fusion Strategies.
Feature fusion is central to multimodal learning, determining how information from distinct modalities is integrated. Based on the reviewed literature, fusion strategies are broadly classified into 3 main categories: early, intermediate, and late fusion, depending on when information from multiple modalities is integrated (57) (Figure 2B).
Early fusion (37 studies) involves merging raw or preprocessed features from multiple modalities before the learning phase, enabling joint representation learning from the outset. For example, several neuroimaging-based studies concatenated atlas-derived functional connectivity measures with structured clinical variables into a single feature vector that was subsequently used as input to a classical ML model or a shallow neural network (34,38,50,58). In such settings, features from different modalities are aligned at the input level, without modality-specific encoders, and the downstream model learns directly from the combined feature space. While this approach promotes holistic modeling, it often struggles with issues related to modality imbalance, data noise, and incompatible feature scales (59,60).
Intermediate fusion (35 studies) follows a modular design in which each modality is first encoded into latent representations using modality-specific architectures and then aggregated for joint modeling. This provides flexibility, allowing each modality to be processed with architectures best suited to its structure, e.g., convolution based for imaging and Transformer based for language (22,55,61,62). A representative example is the use of graph neural networks (GNNs) to encode brain connectivity or symptom co-occurrence structures separately, with learned node or graph embeddings subsequently being combined for prediction, as demonstrated in multiscale adaptive multichannel fusion graph convolutional network (39) and MS2-GNN (20). Similarly, attention-based fusion can dynamically weight and align cross-modal features during learning to enhance context-sensitive integration, e.g., co-attention fusion (63), fusion transformer module (64), and dynamic restrained uncertainty weighting (65).
Late fusion (12 studies) preserves separate modeling pipelines for each modality until the final decision layer, where predictions are aggregated using ensemble, voting, or probability averaging. For example, some studies trained separate classifiers on text, audio, and physiological signals and combined their outputs to generate a final diagnostic or risk score, without enforcing shared feature representations across modalities (19,45,66,67). This strategy is particularly useful when dealing with missing modalities or variable data quality, as it avoids forcing alignment at the feature or representation level and allows independent optimization of each modality-specific model.
AI Architectures in Multimodal Psychiatry Research.
To characterize the methodological landscape of AI applications in multimodal psychiatry, the 88 reviewed studies were categorized into 4 broad groups based on their dominant modeling approach: traditional ML, DL (nontransformer), transformer architectures, and LLMs (Table 2; Figures 2B and 3B). Traditional ML encompasses non-neural, algorithmic methods characterized by fixed model structures and comparatively shallow optimization routines (68,69). DL (nontransformer) refers to neural network–based architectures capable of hierarchical feature learning (70,71). Transformer models, distinguished by their self-attention mechanism, represent a more recent and flexible class of DL architectures. Unlike earlier DL models that process data sequentially or spatially, transformers can model global dependencies in input sequences regardless of position. This makes them particularly powerful for learning from multimodal data such as text, speech, and imaging, explaining their growing role in psychiatric AI research (72).
Table 2.
AI Model Architectures
| Main category | Subcategories | Description | Examples* |
|---|---|---|---|
| Traditional Machine Learning | Linear | Models that assume a linear relationship between input features and output | Linear Regression, Logistic Regression, Linear SVM. |
| Non-linear | Models that can capture non-linear relationships | Decision Trees, Random Forests, XGBoost, K-Nearest Neighbors | |
| Deep Learning (Non-Transformer) | RNN | Recurrent Neural Networks designed for sequential data, good at modeling temporal dependencies | LSTM, GRU |
| CNN | Convolutional Neural Networks specialized for spatial data, such as images, using convolutional layers. | ResNet, ShuffleNet | |
| Autoencoder | Neural networks designed for unsupervised learning to encode and reconstruct input data | Denoising AE (DAE), Variational AE (VAE) | |
| Graph | Neural networks that operate on graph-structured data | Graph Neural Networks (GNNs), Graph Convolutional Networks (GCNs), Graph Attention Networks (GAT) | |
| Hybrid | Models combining different neural network architectures | CNN-RNN, LSTM-ShuffleNet | |
| Other | Other neural network types that do not fit in the above categories | Generative Adversarial Networks (GANS), Spatiotemporal Deep Neural networks (stDNN) | |
| Transformer | Encoder-only | Transformer models that use only the encoder stack | BERT, RoBERTa |
| Other | Transformer architectures (either encoder or decoder) + other DL/ML architectures | Vision transformer (ViT) + Adversarial Networks | |
| LLM | Multi-Modality LLM | LLMs that can process and reason over multiple modalities (e.g. text, image, audio, video). | GPT-4o and later GPT series Gemini (Pro / Flash) Claude 3 |
Study-wise analyses are referred to Supplementary 1.
Traditional ML (25 studies) remains foundational in psychiatry, especially in research relying on structured phenotypic or clinical data. These methods were further subdivided based on their ability to model linearity. Linear models (9 studies), such as logistic regression and linear support vector machines (SVMs), assume additive relationships between features and outcomes, making them suitable for hypothesis-driven modeling or scenarios requiring interpretability (35,50,73). Nonlinear models (10 studies), including random forests and kernel-based SVMs, capture complex, interaction-driven relationships. They have often been used to integrate heterogeneous features such as cognitive, behavioral, and imaging-derived variables (34,58,74,75). Six studies used both linear and nonlinear models in exploratory or comparative settings to assess robustness under different modeling assumptions (23,76,77).
DL (nontransformer) architectures were applied in 39 studies, representing a shift toward data-driven and representation learning paradigms. Among these, GNNs (12 studies) were the most frequently used DL subcategory (61,78,79). GNNs encoded relationships across imaging, behavioral, and clinical data into structured graph representations, with some models incorporating graph attention mechanisms or edge-based learning (37,39,80). Several studies also implemented multiscale or multichannel GNNs to integrate local and global topological features, as seen in frameworks such as multimodal contrastive representation learning network (61) and WL-DeepGCN (62). Convolutional neural networks (CNNs) (4 studies) were used in applications involving image or video data such as structural or functional MRI (fMRI) (22,81–83). Recurrent neural networks (RNNs) appeared in 2 studies involving temporal or sequential data (53,84), but their overall underuse suggests that the modeling of longitudinal psychiatric trajectories remains relatively underdeveloped or has been replaced by more recent temporal architectures. Hybrid architectures (6 studies) combined CNNs, RNNs, or autoencoders to enable hierarchical or cross-modal modeling, especially in studies integrating imaging with behavioral or linguistic data (20,32,33,85). The remaining 11 studies fitting into the Other subcategory used methods such as generative adversarial network (GAN) (42,86,87), reinforcement learning (88,89), and multilayer perceptron (21,90,91).
Transformer-based models (18 studies) constituted the most rapidly growing and novel category due to their ability to model long-range dependencies, handle diverse data types, and support joint representation learning across modalities (92). Across the reviewed studies, encoder-only transformer architectures (13 studies), which were primarily used as flexible feature learners, predominated. Early work concentrated on language-centric applications, where pretrained encoders such as BERT (24) or XLNet (31) were applied to clinical notes or semistructured EHR data, often enabling scalable modeling across multiple psychiatric diagnoses (93,94). More recent studies extended encoder-only transformers to multimodal settings, incorporating behavioral features (95), physiological time series (65,96), and neuroimaging data (64,87,97), indicating a broader trend toward unifying diverse modalities within shared latent representations. A couple of studies explored temporal encoder-decoder transformer variants for temporal monitoring and affective state forecasting using wearable (98) and mobile sensing data (30), highlighting emerging interest in modeling psychiatric dynamics over time. The remaining 3 studies, categorized as Other, integrated transformer components with complementary DL architectures such as GAN (87) and residual graph attention network (99). Combining transformer components with advanced DL architectures enables capturing complex spatial and cross-domain structure in multimodal data.
Among the included studies, only 2 used LLMs as the primary predictive architecture. In the first study (100), GPT-4o was used to interpret multimodal digital phenotyping data by jointly processing textual representations of active self-report and passive sensing features. In the second study, Social-RecNet (101) integrated speech acoustics and dialogue transcripts through cross-attention–based alignment and representation-level fusion before projecting them into an LLM embedding space for Autism Diagnostic Observation Schedule score prediction. Although multimodal LLMs (MLLMs) have rapidly emerged and are increasingly being explored for multimodal clinical question answering and decision support, their use as the primary architecture for quantitative predictive modeling in psychiatry research remains rare (102–105).
Psychiatry Applications
From the perspective of multimodal applications in psychiatry, 88 selected studies were organized into 2 overarching categories: 1) present-state characterization, focusing on determining an individual’s current mental health status, including disease identification, diagnosis, subtyping, and patient phenotyping; and 2) future-state prediction, which aims to forecast clinical trajectories, including onset risk, disease progression, treatment response, and long-term outcomes (Table 3; Figures 2C and 3C).
Table 3.
Psychiatry Applications
| Categories | Subcategories | Description | Examples |
|---|---|---|---|
| Present-State Characterization | Disease Identification and Diagnosis | Recognizing whether a disease is present and assigning the correct disease label, including differentiating among similar conditions. |
|
| Disease Characterization and Subtyping | Organizing a disease into subtypes, stages, or other clinically meaningful labels, such as severity, location, or etiology. |
|
|
| Patient Phenotyping and Stratification | Grouping patients based on shared clinical, genetic, or behavioral features to identify subpopulations that may benefit from different management strategies. |
|
|
| Future-State Prediction | Risk and Onset Prediction | Estimating the likelihood or timing of a disease or symptom developing in the future. |
|
| Prognosis and Progression Prediction | Forecasting the future course, severity, or progression trajectory of an existing condition. |
|
|
| Treatment Response and Outcome Prediction | Estimating the effectiveness, tolerability, or outcome of a specific intervention or treatment. |
|
Present-State Characterization.
Disease identification and diagnosis accounted for 51 studies, covering major depressive disorder (MDD), schizophrenia, autism spectrum disorder (ASD), and attention-deficit/hyperactivity disorder (ADHD), among others. Multimodal data integration consistently improved diagnostic accuracy relative to unimodal models, addressing longstanding challenges of misdiagnosis and underdiagnosis. High discriminative performance was frequently reported. For example, in MDD, a multimodal model achieved an area under the curve (AUC) of 0.94 for younger adult depression (91), while another DL model integrating structural MRI and resting-state fMRI yielded 75% accuracy (63). In schizophrenia, functional connectivity approaches achieved 98% accuracy (106), enabling earlier recognition and intervention. In ASD, ensemble DL models reached 96% accuracy (77).
Disease subtyping and characterization (13 studies) aims to identify clinically and biologically meaningful subgroups, or salient disease features, such as severity and underlying mechanisms. Unlike tasks with manually defined labels, disease subtyping and characterization are often data driven, using unsupervised learning to discover latent structures in multimodal data. The resulting clusters or subgroup representations can subsequently be incorporated as informative features or labels for training downstream AI models. For example, Zhu et al. (107) utilized functional connectivity features derived from MRI to partition individuals with alcohol and nicotine use disorders into distinct biotypes, which were then used to enhance a multitask classification framework (108). Similarly, Raschka et al. (109) clustered longitudinal disease trajectories in Huntington disease and identified 2 progression subtypes, which subsequently served as labels for training an XGBoost classifier to predict disease progression. In addition to clustering-based approaches, Rodríguez-Herrera et al. (88) utilized reinforcement learning to identify contingency-based flexibility mechanisms for adults with ADHD.
Patient phenotyping and stratification studies (13 studies) targeted within-disorder variability, with an emphasis on identifying clinically actionable patient subgroups. In populations at risk of delirium, multimodal ML models reported AUCs of up to 0.94 (110), enabling actionable risk stratification for early intervention and preventive care. Unlike purely data-driven subtyping, phenotyping and stratification approaches typically align learned patient representations with clinically meaningful outcomes, thereby facilitating risk assessment and decision support in real-world settings.
Future-State Prediction.
Risk and onset prediction (18 studies) investigated the likelihood of developing psychiatric disorders prior to full clinical manifestation (29,50,75,89). Models integrating genetics, clinical, and psychosocial features achieved robust performance, e.g., an AUC of 0.79 in adolescent depression onset (111) and 91% accuracy for developmental disorders (45). Importantly, these models operate explicitly along a temporal dimension, aiming to identify high-risk individuals before disease onset.
Treatment response and outcome prediction were addressed in 16 studies, with MDD as the most studied disorder. Multimodal models integrating clinical, genetic, and imaging features predicted antidepressant efficacy with 83% specificity (112). Outcome prediction for suicide attempts reached an AUC of 0.80 (28). In schizophrenia, repetitive transcranial magnetic stimulation response prediction achieved balanced accuracy of 80% (113).
Nine studies focused on predicting the course and progression of psychiatric disorders over time. A nationwide study integrating registry, family history, and genetic data predicted severe progression trajectories with an AUC of 0.72 (21), providing stratification relevant for resource allocation. In treatment-resistant depression, dynamic cingulate activity tracked with deep brain stimulation highlights neurophysiological markers for recovery (54).
In the overall quantitative analysis, MDD emerged as the most frequently studied disorder, followed by schizophrenia, autism, and bipolar disorder (Figure 3D). Disease identification and diagnosis remained the dominant application area. As illustrated in Figure 5, 6 primary psychiatric disorders are paired to 3 analytic dimensions—data modalities, AI architectures, and applications. For example, studies on MDD primarily utilized text, signal, and imaging data, most often with DL models for disease identification and diagnosis. Similar trends have been observed across other disorders, where disease diagnosis remains central. Notably, this mapping also revealed underexplored modality-architecture-task combinations, thereby providing a road map for future explorations.
Figure 5.

Connection of primary psychiatric disorders with data modalities, artificial intelligence (AI) architectures, and psychiatry applications. (A–F) For each psychiatric disorder, data modalities, application types, and AI architectures are arranged in a clockwise order, ranked by the study numbers in a descending way. Numbers on the edges connecting disorder nodes (red) and the 3-dimensional nodes (purple, blue, or green) indicate the number of studies. An individual study can have multiple subcategories within each dimension, particularly for data modality (multimodality) and psychiatry application (multitask). ADHD, attention-deficit/hyperactivity disorder; ANX, anxiety disorder; ASD, autism spectrum disorder; BD, bipolar disorder; DL, deep learning; G&M, genetic and molecular; LLM, large language model; MDD, major depressive disorder; ML, machine learning; SCZ, schizophrenia.
DISCUSSION
Challenges
Despite growing enthusiasm for multimodal AI in psychiatry, a closer examination reveals unresolved conceptual and methodological challenges. These challenges can be broadly categorized into data-related and model-related barriers, each constraining clinical translation in distinct but interconnected ways.
Data-Related Challenges
A central barrier to advancing multimodal AI in psychiatry lies in the limitations of available datasets, which often lack the scale, diversity, and quality needed for robust modeling. Insufficient scale is pervasive. Many cohorts remain small, promoting overfitting and hindering generalizability (19,20). High annotation costs further constrain the development and availability of large, well-curated datasets. One bottleneck is the scarcity of paired and linked modalities. Multisite variability in protocols and equipment (78), as well as inconsistent preprocessing, reduced modal accuracy and generalizability (48,114). Population biases, such as skewed distributions by age (19), sex (62), or restricted demographics (89), constrain the utility of models across diverse groups. Reliance on single datasets and the absence of external validation (80) amplify risks of study-specific confounds. For example, among the studies reviewed, only 11 of 88 (12.5%) conducted external validation using independent datasets, multisite cohorts, or cross-hospital testing.
An additional and increasingly salient data-related challenge is the prevalence of missing modalities. Multimodal models trained under fully paired assumptions can degrade substantially when modalities are absent (115–117). Prior work has attempted to mitigate this issue through strategies such as decision-level fusion (118), modality imputation (119–122), and shared-representation or alignment-based learning. However, these approaches limit cross-modal interaction (123–125), rely on strong prior assumptions, or struggle when modalities exhibit high structural heterogeneity (126). This limitation is especially relevant in biological psychiatry, where neuroimaging, physiological signals, behavioral data, and clinical text differ markedly in availability, scale, and noise characteristics.
Single point–based proxy labels and single scale–based input data also pose significant limitations. Specifically, many models are trained on proxy labels derived from self-report instruments. These tools capture symptom severity at a single time point and are highly sensitive to mood fluctuations or transient stressors. Using such instruments as ground truth necessitates rigorous validation, particularly through longitudinal follow-up to ensure the predictions correspond to stable diagnostic outcomes.
Model-Related Challenges
Despite rapid advances in AI architectures, several challenges hinder the reliable application of multimodal models in psychiatry. Generalizability remains a concern, as domain shifts between pretrained models and psychiatric datasets often lead to suboptimal performance in real-world settings. Fusion of heterogeneous modalities is technically difficult due to divergent temporal and spatial properties, often requiring handcrafted or heuristic-based fusion strategies that may not scale well (24,55). Interpretability is frequently under-addressed, especially in deep architectures such as transformers and GANs (51). Robustness to missing or noisy data is another weakness (24,127). Finally, although ensemble and hybrid models have shown improved accuracy, their computational burden limits their deployment in resource-constrained clinical environments (31,87).
Opportunities
The rapid growth of AI-driven multimodal learning also opens important avenues for innovation and opportunities in psychiatry, with emerging roles of MLLMs at the methodological frontier. The integration of LLMs into psychiatric research is in its early stages. Recent efforts, such as psychiatric LLM systems [e.g., PsyLLM (128), MAGI (129)] and emerging medical MLLMs (104,130), reflect growing interest but remain largely text centric. These findings underscore a substantial opportunity; contemporary MLLM architectures are underutilized and hold a substantial potential to advance psychiatric research and clinical practice.
Improving model generalizability remains a critical goal. Domain adaptation and federated learning are promising strategies to reduce domain shifts and extend modal applicability (22,24,30,45,55). In parallel, innovations in learning paradigms, including self-supervised, contrastive, and multitask learning, promote information sharing across modalities (81–83). Interpretability can be enhanced through attention visualization, feature attribution, and ablation studies, which reveal clinically relevant drivers of predictions, thereby enhancing model transparency (51,62,78,79,114).
Model innovation should be paired with progress in data quality and access. Development of large, diverse, and modality-paired datasets is essential to accelerate both training and validation of multimodal models. Multisite, multi-disciplinary collaborations are especially promising, enabling harmonization, normalization, and external validation across heterogeneous cohorts (24,87,131). The Individually Measured Phenotypes to Advance Computational Translation in Mental Health (IMPACT-MH) consortium (131) exemplifies such efforts. Through its data coordination center, IMPACT-MH facilitates multimodal data collection, curation, and harmonization, creating patient-level, well-paired resources with strong potential to train psychiatry-focused MLLMs.
Applications of multimodal AI in psychiatry extend across the translational continuum. Automated tools can reduce clinician burden, as illustrated by systems that assess schizophrenia symptom severity from interview recordings, thereby freeing clinicians to focus on patient care (55). Multi-modal approaches also support individualized prognostic stratification and treatment tailoring, thereby advancing the goals of precision psychiatry. Furthermore, they facilitate the discovery of novel psychiatric biomarkers, spanning neuroimaging, cerebrospinal fluid, genomics, digital behavior, and physical activity (29,64,114).
CONCLUSIONS
Multimodal AI represents a transformative frontier in psychiatry research, with the potential to drive the field toward a truly integrative science of mental health. By synthesizing diverse data modalities, AI models can capture the complex interplay of biological, cognitive, and behavioral dimensions that underlie psychiatric disorders, thereby enhancing diagnostic precision, prognostic accuracy, and therapeutic personalization. This review has outlined the breadth of primary data modalities, methodological paradigms, and representative applications that reveal the field’s rapid momentum. However, unresolved data and model-related challenges continue to impede real-world applicability. Overcoming these barriers will require sustained methodological innovations, the generation of richer, better-linked, and more representative datasets, and rigorous validation across sites and populations. Deep collaboration across psychiatry, neuroscience, computer science, and informatics is a necessity to bridge disciplinary boundaries, to ensure that technological innovation remains anchored in clinical needs, and for AI-driven multimodal learning can move from proof-of-concept toward delivering tangible real-world impact. Ultimately, the convergence of multimodal data, AI innovation, and clinical insight provides a foundation for a new era of precision psychiatry in which diagnosis, prognosis, and treatment are tailored to the multidimensional realities of each patient.
Supplementary Material
Supplementary material cited in this article is available online at https://doi.org/10.1016/j.bpsc.2026.03.013.
ACKNOWLEDGMENTS AND DISCLOSURES
This work was supported by the National Institute of Mental Health (Grant No. U24MH136069 [to HX, YC, and CT]) and the National Institute on Aging (Grant Nos. U01AG088076 [to CT] and R01AG072799 [to HX and CT]).
The authors report no biomedical financial interests or potential conflicts of interest.
REFERENCES
- 1.Kieling C, Buchweitz C, Caye A, Silvani J, Ameis SH, Brunoni AR, et al. (2024): Worldwide prevalence and disability from mental disorders across childhood and adolescence: Evidence from the global burden of disease study. JAMA Psychiatry 81:347–356. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.GBD (2019): Mental Disorders Collaborators (2022): Global, regional, and national burden of 12 mental disorders in 204 countries and territories, 1990–2019: A systematic analysis for the Global Burden of Disease Study 2019. Lancet Psychiatry 9:137–150. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.National Institute of Mental Health (NIMH) (2026): Mental illness. Available at: https://www.nimh.nih.gov/health/statistics/mental-illness. Accessed August 20, 2025.
- 4.Columbia Business School (2024): The $282 billion toll: Quantifying the economic impact of mental illness. Available at: https://business.columbia.edu/research-brief/economic-impact-mental-illness. Accessed August 20, 2025.
- 5.Andreassen OA, Hindley GFL, Frei O, Smeland OB (2023): New insights from the last decade of research in psychiatric genetics: Discoveries, challenges and clinical implications. World Psychiatry 22:4–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Segal A, Parkes L, Aquino K, Kia SM, Wolfers T, Franke B, et al. (2023): Regional, circuit and network heterogeneity of brain abnormalities in psychiatric disorders. Nat Neurosci 26:1613–1629. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Wendt FR, Pathak GA, Tylee DS, Goswami A, Polimanti R (2020): Heterogeneity and Polygenicity in psychiatric disorders: A genome-wide perspective. Chronic Stress (Thousand Oaks) 4: 2470547020924844. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Porter A, Fei S, Damme KSF, Nusslock R, Gratton C, Mittal VA (2023): A meta-analysis and systematic review of single vs. multi-modal neuroimaging techniques in the classification of psychosis. Mol Psychiatry 28:3278–3292. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Zhang Y, Wang J, Zong H, Singla RK, Ullah A, Liu X, et al. (2025): The comprehensive clinical benefits of digital phenotyping: From broad adoption to full impact. NPJ Digit Med 8:196. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Sathyanarayanan A, Mueller TT, Ali Moni M, Schueler K, ECNP TWG Network members, Baune BT, et al. (2023): Multi-omics data integration methods and their applications in psychiatric disorders. Eur Neuropsychopharmacol 69:26–46. [DOI] [PubMed] [Google Scholar]
- 11.Jiao Y, Zhao K, Wei X, Carlisle NB, Keller CJ, Oathes DJ, et al. (2025): Deep graph learning of multimodal brain networks defines treatment-predictive signatures in major depression. Mol Psychiatry 30:3963–3974. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.ChatGPT (2026): ChatGPT. Available at: https://chatgpt.com/?locale=en-US. Accessed August 22, 2025.
- 13.Hua Y, Na H, Li Z, Liu F, Fang X, Clifton D, Torous J (2025): A scoping review of large language models for generative tasks in mental health care. NPJ Digit Med 8:230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Perlis RH, Goldberg JF, Ostacher MJ, Schneck CD (2024): Clinical decision support for bipolar depression using large language models. Neuropsychopharmacology 49:1412–1416. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Kallstenius T, Capusan AJ, Andersson G, Williamson A (2025): Comparing traditional natural language processing and large language models for mental health status classification: A multi-model evaluation. Sci Rep 15:24102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Sha Y, Pan H, Xu W, Meng W, Luo G, Du X, et al. (2025): MDD-LLM: Towards accuracy large language models for major depressive disorder diagnosis. J Affect Disord 388:119774. [DOI] [PubMed] [Google Scholar]
- 17.Liu JJ, Borsari B, Li Y, Liu SX, Gao Y, Xin X, et al. (2025): Digital phenotyping from wearables using AI characterizes psychiatric disorders and identifies genetic associations. Cell 188:515–529.e15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Webb CA, Ren B, Rahimi-Eichi H, Gillis BW, Chung Y, Baker JT (2025): Personalized prediction of negative affect in individuals with serious mental illness followed using long-term multimodal mobile phenotyping. Transl Psychiatry 15:174. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Nash C, Nair R, Naqvi SM (2024): Insights into detecting adult ADHD symptoms through advanced dual-stream machine learning. IEEE Trans Neural Syst Rehabil Eng 32:3378–3387. [DOI] [PubMed] [Google Scholar]
- 20.Chen T, Hong R, Guo Y, Hao S, Hu B (2023): MS2-GNN: Exploring GNN-based multimodal fusion network for depression detection. IEEE Trans Cybern 53:7749–7759. [DOI] [PubMed] [Google Scholar]
- 21.Allesøe RL, Thompson WK, Bybjerg-Grauholm J, Hougaard DM, Nordentoft M, Werge T, et al. (2023): Deep learning for cross-diagnostic prediction of mental disorder diagnosis and prognosis using Danish nationwide register and genetic data. JAMA Psychiatry 80:146–155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Othmani A, Zeghina A-O, Muzammel M (2022): A model of normality inspired deep learning framework for depression relapse prediction using audiovisual data. Comput Methods Programs Biomed 226:107132. [DOI] [PubMed] [Google Scholar]
- 23.Ho CSH, Wang J, Tay GWN, Ho R, Lin H, Li Z, Chen N (2025): Application of functional near-infrared spectroscopy and machine learning to predict treatment response after six months in major depressive disorder. Transl Psychiatry 15:7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Yang H, Zhu D, He S, Xu Z, Liu Z, Zhang W, Cai J (2024): Enhancing psychiatric rehabilitation outcomes through a multimodal multitask learning model based on BERT and TabNet: An approach for personalized treatment and improved decision-making. Psychiatry Res 336:115896. [DOI] [PubMed] [Google Scholar]
- 25.Lens.org (2026): The lens—Patent and scholarly search and analysis. Available at: https://www.lens.org/lens. Accessed January 14, 2026.
- 26.Meta (2026): The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation. Available at: https://ai.meta.com/blog/llama-4-multimodal-intelligence/. Accessed August 11, 2025.
- 27.Yang A, Li A, Yang B, Zhang B, Hui B, Zheng B, et al. (2025): Qwen3 technical report. arXiv 10.48550/arXiv.2505.09388. [DOI] [Google Scholar]
- 28.Zheng S, Zeng W, Wu Q, Li W, He Z, Li E, et al. (2024): Predictive models for suicide attempts in major depressive disorder and the contribution of EPHX2: A pilot integrative machine learning study. Depress Anxiety 2024:5538257. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Yokotani K, Takano M, Abe N, Kato TA (2025): Predicting social anxiety disorder based on communication logs and social network data from a massively multiplayer online game: Using a graph neural network. Psychiatry Clin Neurosci 79:274–281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Paz-Arbaizar L, Lopez-Castroman J, Artés-Rodríguez A, Olmos PM, Ramírez D (2025): Emotion forecasting: A transformer-based approach. J Med Internet Res 27:e63962. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Zhu E, Wang J, Zhou G, Li C, Chen F, Ju K, et al. (2025): A highly scalable deep learning language model for common risks prediction among psychiatric inpatients. BMC Med 23:308. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Hong M, Kang R-R, Yang JH, Rhee SJ, Lee H, Kim Y-G, et al. (2024): Comprehensive symptom prediction in inpatients with acute psychiatric disorders using wearable-based deep learning models: Development and validation study. J Med Internet Res 26:e65994. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Allesøe RL, Nudel R, Thompson WK, Wang Y, Nordentoft M, Børglum AD, et al. (2022): Deep learning-based integration of genetics with registry data for stratification of schizophrenia and depression. Sci Adv 8:eabi7293. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Lee DY, Byeon G, Kim N, Son SJ, Park RW, Park B (2024): Neuroimaging and natural language processing-based classification of suicidal thoughts in major depressive disorder. Transl Psychiatry 14:276. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Krakowski K, Oliver D, Arribas M, Stahl D, Fusar-Poli P (2024): Dynamic and transdiagnostic risk calculator based on natural language processing for the prediction of psychosis in secondary mental health care: Development and internal-external validation cohort study. Biol Psychiatry 96:604–614. [DOI] [PubMed] [Google Scholar]
- 36.Hansen L, Bernstorff M, Enevoldsen K, Kolding S, Damgaard JG, Perfalk E, et al. (2025): Predicting diagnostic progression to schizophrenia or bipolar disorder via machine learning. JAMA Psychiatry 82:459–469. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Chen Y, Yan J, Jiang M, Zhang T, Zhao Z, Zhao W, et al. (2024): Adversarial learning based node-edge graph attention networks for autism spectrum disorder identification. IEEE Trans Neural Netw Learn Syst 35:7275–7286. [DOI] [PubMed] [Google Scholar]
- 38.Zhu C, Fu Z, Chen L, Yu F, Zhang J, Zhang Y, et al. (2022): Multimodality connectome-based predictive modeling of individualized compulsions in obsessive-compulsive disorder. J Affect Disord 311:595–603. [DOI] [PubMed] [Google Scholar]
- 39.Pan J, Lin H, Dong Y, Wang Y, Ji Y (2022): MAMF-GCN: Multi-scale adaptive multi-channel fusion deep graph convolutional network for predicting mental disorder. Comput Biol Med 148:105823. [DOI] [PubMed] [Google Scholar]
- 40.Cameron C, Yassine B, Carlton C, Francois C, Alan E, András J, et al. (2013): The Neuro Bureau Preprocessing Initiative: Open sharing of preprocessed neuroimaging data and derivatives. Front Neuroinform 7. [Google Scholar]
- 41.Tzourio-Mazoyer N, Landeau B, Papathanassiou D, Crivello F, Etard O, Delcroix N, et al. (2002): Automated anatomical labeling of activations in SPM using a macroscopic anatomical parcellation of the MNI MRI single-subject brain. Neuroimage 15:273–289. [DOI] [PubMed] [Google Scholar]
- 42.Corponi F, Li BM, Anmella G, Mas A, Pacchiarotti I, Valentí M, et al. (2024): Automated mood disorder symptoms monitoring from multivariate time-series sensory data: Getting the full picture beyond a single number. Transl Psychiatry 14:161. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Wiebe A, Selaskowski B, Paskin M, Asché L, Pakos J, Aslan B, et al. (2024): Virtual reality-assisted prediction of adult ADHD based on eye tracking, EEG, actigraphy and behavioral indices: A machine learning analysis of independent training and test samples. Transl Psychiatry 14:508. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Vidivelli S, Padmakumari P, Shanthi P (2025): Multimodal autism detection: Deep hybrid model with improved feature level fusion. Comput Methods Programs Biomed 260:108492. [DOI] [PubMed] [Google Scholar]
- 45.Liu X, Hasan MR, Gedeon T, Hossain MZ (2024): MADE-for-ASD: A multi-atlas deep ensemble network for diagnosing autism spectrum disorder. Comput Biol Med 182:109083. [DOI] [PubMed] [Google Scholar]
- 46.Hamilton M (1960): A rating scale for depression. J Neurol Neurosurg Psychiatry 23:56–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Young RC, Biggs JT, Ziegler VE, Meyer DA (1978): A rating scale for mania: Reliability, validity and sensitivity. Br J Psychiatry 133:429–435. [DOI] [PubMed] [Google Scholar]
- 48.Gao J, Tang H, Wang Z, Li Y, Luo N, Song M, et al. (2025): Graph neural networks and multimodal DTI features for schizophrenia classification: Insights from brain network analysis and gene expression. Neurosci Bull 41:933–950. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Checa-Robles FJ, Salvetat N, Cayzac C, Menhem M, Favier M, Vetter D, et al. (2024): RNA editing signatures powered by artificial intelligence: A new frontier in differentiating schizophrenia, bipolar, and schizoaffective disorders. Int J Mol Sci 25:12981. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Zhang X, Bhatt RR, Todorov S, Gupta A (2023): Brain-gut micro-biome profile of neuroticism predicts food addiction in obesity: A transdiagnostic approach. Prog Neuropsychopharmacol Biol Psychiatry 125:110768. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Dia M, Khodabandelou G, Othmani A (2024): Paying attention to uncertainty: A stochastic multimodal transformers for post-traumatic stress disorder detection using video. Comput Methods Programs Biomed 257:108439. [DOI] [PubMed] [Google Scholar]
- 52.Chen J, Chan NY, Li C-T, Chan JWY, Liu Y, Li SX, et al. (2024): Multimodal digital assessment of depression with actigraphy and app in Hong Kong Chinese. Transl Psychiatry 14:150. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Zhang L, Fan Y, Jiang J, Li Y, Zhang W (2022): Adolescent depression detection model based on multimodal data of interview audio and text. Int J Neur Syst 32:2250045. [DOI] [PubMed] [Google Scholar]
- 54.Alagapan S, Choi KS, Heisig S, Riva-Posse P, Crowell A, Tiruvadi V, et al. (2023): Cingulate dynamics track depression recovery with deep brain stimulation. Nature 622:130–138. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Chuang C-Y, Lin Y-T, Liu C-C, Lee L-E, Chang H-Y, Liu A-S, et al. (2023): Multimodal assessment of schizophrenia symptom severity from linguistic, acoustic and visual cues. IEEE Trans Neural Syst Rehabil Eng 31:3469–3479. [DOI] [PubMed] [Google Scholar]
- 56.Xu X, Li J, Zhu Z, Zhao L, Wang H, Song C, et al. (2024): A comprehensive review on synergy of multi-modal data and AI technologies in medical diagnosis. Bioengineering (Basel) 11:219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Kline A, Wang H, Li Y, Dennis S, Hutch M, Xu Z, et al. (2022): Multimodal machine learning in precision health: A scoping review. NPJ Digit Med 5:171. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Hassan I, Nahid N, Islam M, Hossain S, Schuller B, Ahad MAR (2025): Automated autism assessment with multimodal data and ensemble learning: A scalable and consistent robot-enhanced therapy framework. IEEE Trans Neural Syst Rehabil Eng 33:1191–1201. [DOI] [PubMed] [Google Scholar]
- 59.Jiao T, Guo C, Feng X, Chen Y, Song J (2024): A comprehensive survey on deep learning multi-modal fusion: Methods, technologies and applications. Comput Mater Contin 80:1–35. [Google Scholar]
- 60.Teoh JR, Dong J, Zuo X, Lai KW, Hasikin K, Wu X (2024): Advancing healthcare through multimodal data fusion: A comprehensive review of techniques and applications. PeerJ Comput Sci 10:e2298. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Kong Y, Wang W, Liu X, Gao S, Hou Z, Xie C, et al. (2023): Multi-connectivity representation learning network for major depressive disorder diagnosis. IEEE Trans Med Imaging 42:3012–3024. [DOI] [PubMed] [Google Scholar]
- 62.Wang M, Guo J, Wang Y, Yu M, Guo J (2023): Multimodal autism spectrum disorder diagnosis method based on DeepGCN. IEEE Trans Neural Syst Rehabil Eng 31:3664–3674. [DOI] [PubMed] [Google Scholar]
- 63.Zheng G, Zheng W, Zhang Y, Wang J, Chen M, Wang Y, et al. (2023): An attention-based multi-modal MRI fusion model for major depressive disorder diagnosis. J Neural Eng 20:066005. [DOI] [PubMed] [Google Scholar]
- 64.Wang G, Fan F, Shi S, An S, Cao X, Ge W, et al. (2024): Multi modality fusion transformer with spatio-temporal feature aggregation module for psychiatric disorder diagnosis. Comput Med Imaging Graph 114:102368. [DOI] [PubMed] [Google Scholar]
- 65.Song M, Yang Z, Triantafyllopoulos A, Zhang Z, Nan Z, Tang M, et al. (2024): Empowering mental health monitoring using a macro-micro personalization framework for multimodal-multitask learning: Descriptive study. JMIR Ment Health 11:e59512. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Grover N, Chharia A, Upadhyay R, Longo L (2023): Schizo-Net: A novel schizophrenia diagnosis framework using late fusion multi-modal deep learning on electroencephalogram-based brain connectivity indices. IEEE Trans Neural Syst Rehabil Eng 31:464–473. [DOI] [PubMed] [Google Scholar]
- 67.Chavanne AV, Meinke C, Langhammer T, Roesmann K, Boehnlein J, Gathmann B, et al. (2023): Individual-level prediction of exposure therapy outcome using structural and functional MRI data in spider phobia: A machine-learning study. Depress Anxiety 2023:8594273. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Alaskar H, Saba T (2021): Machine learning and deep learning: A comparative review. In: Singh Mer KK, Semwal VB, Bijalwan V, Crespo RG, editors. Proceedings of the Integrated Intelligence Enable Networks and Computing. Singapore: Springer, 143–150. [Google Scholar]
- 69.Quinn TP, Hess JL, Marshe VS, Barnett MM, Hauschild A-C, Maciukiewicz M, et al. (2024): A primer on the use of machine learning to distil knowledge from data in biological psychiatry. Mol Psychiatry 29:387–401. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Shamshirband S, Fathi M, Dehzangi A, Chronopoulos AT, Alinejad-Rokny H (2021): A review on deep learning approaches in healthcare systems: Taxonomies, challenges, and open issues. J Biomed Inform 113:103627. [DOI] [PubMed] [Google Scholar]
- 71.Alzubaidi L, Zhang J, Humaidi AJ, Al-Dujaili A, Duan Y, Al-Shamma O, et al. (2021): Review of deep learning: Concepts, CNN architectures, challenges, applications, future directions. J Big Data 8:53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. (2023): Attention is all you need. arXiv 10.48550/arXiv.1706.03762. [DOI] [Google Scholar]
- 73.Buciuman M-O, Oeztuerk OF, Popovic D, Enrico P, Ruef A, Bieler N, et al. (2023): Structural and functional brain patterns predict formal thought disorder’s severity and its persistence in recent-onset psychosis: Results from the PRONIA study. Biol Psychiatry Cogn Neurosci Neuroimaging 8:1207–1217. [DOI] [PubMed] [Google Scholar]
- 74.Jang S, Sun TH, Shin S, Lee H-J, Shin Y-B, Yeom JW, et al. (2024): A digital phenotyping dataset for impending panic symptoms: A prospective longitudinal study. Sci Data 11:1264. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Karcher NR, Sotiras A, Niendam TA, Walker EF, Jackson JJ, Barch DM (2024): Examining the most important risk factors for predicting youth persistent and distressing psychotic-like experiences. Biol Psychiatry Cogn Neurosci Neuroimaging 9:939–947. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Deng LR, Harmata GIS, Barsotti EJ, Williams AJ, Christensen GE, Voss MW, et al. (2025): Machine learning with multiple modalities of brain magnetic resonance imaging data to identify the presence of bipolar disorder. J Affect Disord 368:448–460. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Wang M, He K, Zhang L, Xu D, Li X, Wang L, et al. (2025): Assessment of glymphatic function and white matter integrity in children with autism using multi-parametric MRI and machine learning. Eur Radiol 35:1623–1636. [DOI] [PubMed] [Google Scholar]
- 78.Liu J, Yang W, Ma Y, Dong Q, Li Y, Hu B, DIRECT Consortium (2024): Effective hyper-connectivity network construction and learning: Application to major depressive disorder identification. Comput Biol Med 171:108069. [DOI] [PubMed] [Google Scholar]
- 79.Jiang C, Lin B, Ye X, Yu Y, Xu P, Peng C, et al. (2024): Graph convolutional network with attention mechanism improve major depressive depression diagnosis based on plasma biomarkers and neuroimaging data. J Affect Disord 360:336–344. [DOI] [PubMed] [Google Scholar]
- 80.Liu L, Xie J, Chang J, Liu Z, Sun T, Qiao H, et al. (2024): H-Net: Heterogeneous neural network for multi-classification of neuropsychiatric disorders. IEEE J Biomed Health Inform 28:5509–5518. [DOI] [PubMed] [Google Scholar]
- 81.Abbas SQ, Chi L, Chen Y-PP (2023): DeepMNF: Deep multimodal neuroimaging framework for diagnosing autism spectrum disorder. Artif Intell Med 136:102475. [DOI] [PubMed] [Google Scholar]
- 82.Dang R, Wang Y, Zhu F, Wang X, Zhao J, Shao P, et al. (2025): Classification of schizophrenia based on RAnet-ET: Resnet based attention network for eye-tracking. J Neural Eng 22:026053. [DOI] [PubMed] [Google Scholar]
- 83.Silva B, Santos L, Barata C, Geminiani A, Fassina G, Rita Gonzalez A, et al. (2024): Attention analysis in robotic-assistive therapy for children with autism. IEEE Trans Neural Syst Rehabil Eng 32:2220–2229. [DOI] [PubMed] [Google Scholar]
- 84.Zantvoort K, Scharfenberger J, Boß L, Lehr D, Funk B (2023): Finding the best match - A case study on the (text-)feature and model choice in digital mental health interventions. J Healthc Inform Res 7:447–479. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Wen G, Cao P, Liu L, Yang J, Zhang X, Wang F, Zaiane OR (2023): Graph self-supervised learning with application to brain networks analysis. IEEE J Biomed Health Inform 27:4154–4165. [DOI] [PubMed] [Google Scholar]
- 86.Liu R, Huang Z-A, Hu Y, Zhu Z, Wong K-C, Tan KC (2024): Attention-like multimodality fusion with data augmentation for diagnosis of mental disorders using MRI. IEEE Trans Neural Netw Learn Syst 35:7627–7641. [DOI] [PubMed] [Google Scholar]
- 87.Bi Y, Abrol A, Jia S, Sui J, Calhoun VD (2024): Gray matters: ViT-GAN framework for identifying schizophrenia biomarkers linking structural MRI and functional network connectivity. Neuroimage 297:120674. [DOI] [PubMed] [Google Scholar]
- 88.Rodríguez-Herrera R, León JJ, Fernández-Martín P, Sánchez-Kuhn A, Soto-Ontoso M, Amaya-Pascasio L, et al. (2025): Contingency-based flexibility mechanisms through a reinforcement learning model in adults with attention-deficit/hyperactivity disorder and obsessive–compulsive disorder. Compr Psychiatry 139:152589. [DOI] [PubMed] [Google Scholar]
- 89.Kaske EA, Chen CS, Meyer C, Yang F, Ebitz B, Grissom N, et al. (2023): Prolonged Physiological Stress Is Associated With a lower rate of Exploratory Learning That is compounded by Depression. Biol Psychiatry Cogn Neurosci Neuroimaging 8:703–711. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Chan L, Simmons C, Tillem S, Conley M, Brazil IA, Baskin-Sommers A (2023): Classifying conduct disorder using a bio-psychosocial model and machine learning method. Biol Psychiatry Cogn Neurosci Neuroimaging 8:599–608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.De Lacy N, Ramshaw MJ, McCauley E, Kerr KF, Kaufman J, Nathan Kutz J (2023): Predicting individual cases of major adolescent psychiatric conditions with artificial intelligence. Transl Psychiatry 13:314. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Cong S, Wang H, Zhou Y, Wang Z, Yao X, Yang C (2024): Comprehensive review of Transformer-based models in neuroscience, neurology, and psychiatry. Brain-X 2:e57. [Google Scholar]
- 93.Huang Y-J, Lin Y-T, Liu C-C, Lee L-E, Hung S-H, Lo J-K, Fu L-C (2022): Assessing schizophrenia patients through linguistic and acoustic features using deep learning techniques. IEEE Trans Neural Syst Rehabil Eng 30:947–956. [DOI] [PubMed] [Google Scholar]
- 94.Yang B, Cao M, Zhu X, Wang S, Yang C, Ni R, Liu X (2024): MMPF: Multimodal purification fusion for automatic depression detection. IEEE Trans Comput Soc Syst 11:7421–7434. [Google Scholar]
- 95.Parra F, Benezeth Y, Yang F (2022): Automatic assessment of emotion dysregulation in American, French, and Tunisian adults and new developments in deep multimodal fusion: Cross-sectional study. JMIR Ment Health 9:e34333. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Li C, Mao Y, Huang Q, Xie W, He X, Wu J (2024): A real-time emotion-aware system based on wireless body area network for IoMT applications. IEEE 11:41182–41193. [Google Scholar]
- 97.Fan Z, Su J, Xu J, Zeng L-L, Hu D (2024): A multibranch cross-attention transformer model for psychiatric disorders diagnosis using multi-modal MRI data. Fuzhoo, China: Presented at the 8th Asian Conference on Artificial Intelligence Technology (ACAIT), November, 8–11. [Google Scholar]
- 98.Kang R-R, Kim Y-G, Hong M, Min Ahn Y, Lee K (2025): AI-based personalized real-time risk prediction for behavioral management in psychiatric wards using multimodal data. Int J Med Inform 198: 105870. [DOI] [PubMed] [Google Scholar]
- 99.Zhu F, Wu B, Huo Y, Dang R, Hu B, Wang Q (2025): MSNet: Multi-modal self-attention network for depression detection via fusion of eye tracking and EEG. In: Proceedings of the 2025 Symposium on Eye Tracking Research and Applications 1–6. New York, NY: ACM, 1–6. [Google Scholar]
- 100.Flathers M, Xia W, Hau C, Nelson BW, Cheong J, Burns J, Torous J (2025): Interpreting psychiatric digital phenotyping data with large language models: A preliminary analysis. BMJ Ment Health 28: e301817. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Chen X-Y, Chen Y-M, Chen C-P, Su B-H, Gau SS-F, Lee C-C (2025): SocialRecNet: A multimodal LLM-based framework for assessing social reciprocity in autism spectrum disorder. In: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India: IEEE, 1–5. [Google Scholar]
- 102.Han T, Adams LC, Bressem KK, Busch F, Nebelung S, Truhn D (2024): Comparative analysis of multimodal large language model performance on clinical vignette questions. JAMA 331:1320–1321. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Moor M, Huang Q, Wu S, Yasunaga M, Zakka C, Dalmia Y, et al. (2023): Med-Flamingo: A multimodal medical few-shot learner. arXiv 10.48550/arXiv.2307.15189. [DOI] [Google Scholar]
- 104.Sellergren A, Kazemzadeh S, Jaroensri T, Kiraly A, Traverse M, Kohlberger T, et al. (2025): MedGemma technical report. arXiv 10.48550/arXiv.2507.05201. [DOI] [Google Scholar]
- 105.AlSaad R, Abd-Alrazaq A, Boughorbel S, Ahmed A, Renault MA, Damseh R, Sheikh J (2024): Multimodal large language models in health care: Applications, challenges, and future outlook. J Med Internet Res 26:e59505. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Alves CL, Toutain TGLO, Porto JAM, Aguiar PMC, De Sena EP, Rodrigues FA, et al. (2023): Analysis of functional connectivity using machine learning and deep learning in different data modalities from individuals with schizophrenia. J Neural Eng 20:056025. [DOI] [PubMed] [Google Scholar]
- 107.Zhu T, Wang W, Chen Y, Kranzler HR, Li C-SR, Bi J (2024): Machine learning of functional connectivity to biotype alcohol and nicotine use disorders. Biol Psychiatry Cogn Neurosci Neuroimaging 9:326–336. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Yang S, Lan Q, Zhang L, Zhang K, Tang G, Huang H, et al. (2025): Multimodal cross-scale context clusters for classification of mental disorders using functional and structural MRI. Neural Netw 185: 107209. [DOI] [PubMed] [Google Scholar]
- 109.Raschka T, Li Z, Gaßner H, Kohl Z, Jukic J, Marxreiter F, Fröhlich H (2024): Unraveling progression subtypes in people with Huntington’s disease. EPMA J 15:275–287. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Friedman JI, Parchure P, Cheng F-Y, Fu W, Cheertirala S, Timsina P, et al. (2025): Machine learning multimodal model for delirium risk stratification. JAMA Netw Open 8:e258874. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Hou J, Mortel L, Popma A, Smit D, van Wingen G (2025): Predicting the onset of mental health problems in adolescents. Psychol Med 55:e128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Bian A, Xiao F, Kong X, Ji X, Fang S, He J, et al. (2024): Predictive modeling of antidepressant efficacy based on cognitive neuropsy-chological theory. J Affect Disord 354:563–573. [DOI] [PubMed] [Google Scholar]
- 113.Dong MS, Rokicki J, Dwyer D, Papiol S, Streit F, Rietschel M, et al. (2024): Multimodal workflows optimally predict response to repetitive transcranial magnetic stimulation in patients with schizophrenia: A multisite machine learning analysis. Transl Psychiatry 14:196. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Gao J, Qian M, Wang Z, Li Y, Luo N, Xie S, et al. (2024): Exploring schizophrenia classification through multimodal MRI and deep graph neural networks: Unveiling brain region-specific weight discrepancies and their association with cell-type specific transcriptomic features. Schizophr Bull 51:217–235. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Ma M, Ren J, Zhao L, Testuggine D, Peng X (2022): Are multimodal transformers robust to missing modality? arXiv 10.48550/arXiv.2204.05454. [DOI] [Google Scholar]
- 116.Huang S-C, Pareek A, Seyyedi S, Banerjee I, Lungren MP (2020): Fusion of medical imaging and electronic health records using deep learning: A systematic review and implementation guidelines. NPJ Digit Med 3:136. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 117.Wang M, Fan S, Li Y, Xie Z, Chen H (2025): Missing-modality enabled multi-modal fusion architecture for medical data. J Biomed Inform 164:104796. [DOI] [PubMed] [Google Scholar]
- 118.Yoo Y, Tang LYW, Li DKB, Metz L, Kolind S, Traboulsee AL, Tam RC (2019): Deep learning of brain lesion patterns and user-defined clinical and MRI features for predicting conversion to multiple sclerosis from clinically isolated syndrome. Comput Methods Biomech Biomed Eng Imaging Vis 7:250–259. [Google Scholar]
- 119.Pan Y, Liu M, Xia Y, Shen D (2022): Disease-image-specific learning for diagnosis-oriented neuroimage synthesis with incomplete multimodality data. IEEE Trans Pattern Anal Mach Intell 44:6839–6853. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.Krysiak-Baltyn K, Nordahl Petersen T, Audouze K, Jørgensen N, Ängquist L, Brunak S (2014): Compass: A hybrid method for clinical and biobank data mining. J Biomed Inform 47:160–170. [DOI] [PubMed] [Google Scholar]
- 121.Ma M, Ren J, Zhao L, Tulyakov S, Wu C, Peng X (2021): SMIL: Multimodal learning with severely missing modality. AAAI 35:2302–2310. [Google Scholar]
- 122.Zhi Z, Liu Z, Elbadawi M, Daneshmend A, Orlu M, Basit A, et al. (2025): Borrowing treasures from neighbors: In-context learning for multimodal learning with missing modalities and data scarcity. Neurocomputing 647:130502. [Google Scholar]
- 123.Wang H, Chen Y, Ma C, Avery J, Hull L, Carneiro G (2023): Multi-modal learning with missing modality via shared-specific feature modelling. arXiv 10.48550/arXiv.2307.14126. [DOI] [Google Scholar]
- 124.Dorent R, Joutard S, Modat M, Ourselin S, Vercauteren T (2019): Hetero-modal variational encoder-decoder for joint modality completion and segmentation. In: Shen D, Liu T, Peters TM, Staib LH, Essert C, Zhou S, editors. (2019), Medical Image Computing and Computer Assisted Intervention—MICCAI 2019, 11765. Cham, Switzerland: Springer International Publishing, 74–82. [Google Scholar]
- 125.Shen Y, Gao M (2019): Brain tumor segmentation on MRI with missing modalities. In: Chung ACS, Gee JC, Yushkevich PA, Bao S, editors. (2019), Information Processing in Medical Imaging, 11492: Cham, Switzerland: Springer International Publishing, 417–428. [Google Scholar]
- 126.Yao W, Yin K, Cheung WK, Liu J, Qin J (2024): DrFuse: Learning disentangled representation for clinical multi-modal fusion with missing modality and modal inconsistency. AAAI 38:16416–16424. [Google Scholar]
- 127.Chen W, Yang J, Sun Z, Zhang X, Tao G, Ding Y, et al. (2024): DeepASD: A deep adversarial-regularized graph learning method for ASD diagnosis with multimodal data. Transl Psychiatry 14:375. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Hu H, Zhou Y, Si J, Wang Q, Zhang H, Ren F, et al. (2025): Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling. arXiv 10.48550/arXiv.2505.15715. [DOI] [Google Scholar]
- 129.Bi G, Chen Z, Liu Z, Wang H, Xiao X, Xie Y, et al. (2025): MAGI: Multiagent guided interview for psychiatric assessment. arXiv 10.48550/arXiv.2504.18260. [DOI] [Google Scholar]
- 130.Team L, Xu W, Chan HP, Li L, Aljunied M, Yuan R, et al. (2025): Lingshu: A Generalist foundation model for unified multimodal medical understanding and reasoning. arXiv 10.48550/arXiv.2506.07044. [DOI] [Google Scholar]
- 131.IMPACT Mental Health (2026): Individually measured phenotypes to advance computational translation in mental health (IMPACT-MH). Available at: https://impact-mh.org/. Accessed September 14, 2025.
- 132.National Library of Medicine (2026): MeSH Browser. Available at: https://meshb.nlm.nih.gov/. Accessed January 19, 2026.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
