Skip to main content
Psychiatric Research and Clinical Practice logoLink to Psychiatric Research and Clinical Practice
. 2026 Aug 11:10.1176/appi.prcp.20260059. Online ahead of print. doi: 10.1176/appi.prcp.20260059

Artificial Intelligence in Psychiatry: Five Decades of Progress and Persistent Translational Challenges

Esteban Zavaleta‐Monestel 1,✉, Luis Guillermo Herrera‐Jiménez 2, Sofía Suárez‐Sánchez 1, Sebastián Arguedas‐Chacón 1, Jeaustin Mora‐Jiménez 1, Ricardo Millán‐González 3
PMCID: PMC13458306  PMID: 42582721

Abstract

Objective

This review examined the historical development of artificial intelligence (AI) in psychiatry from 1972 to 2025 and identified persistent barriers to clinical translation.

Method

The review used a structured source‐identification approach across PubMed/MEDLINE, Web of Science, Google Scholar, ScienceDirect, citation tracking, and targeted historical searches. It synthesized sources qualitatively and organized them chronologically. A structured evidence table summarized representative milestones by era, AI paradigm, clinical task, data modality, validation approach, implementation status, and translational limitation.

Results

Psychiatric AI evolved from symbolic simulation and rule‐based expert systems to connectionist models, supervised machine learning, computational psychiatry, digital phenotyping, multimodal monitoring, digital mental health tools, and large language models. Despite increasing computational sophistication, recurring barriers persisted, including uncertain target validity, diagnostic heterogeneity, limited external validation, poor transportability, interpretability challenges, workflow integration, equity, patient trust, and governance. Across eras, technical progress was cyclical rather than linear, with successive waves reproducing unresolved clinical and implementation challenges.

Conclusions

The clinical impact of psychiatric AI will likely depend less on algorithmic novelty alone than on clearer clinical targets, prospective validation, implementation trials, patient‐centered evaluation, equity‐sensitive generalizability, and mental health–specific governance.

Relevance to Clinical Practice

For routine psychiatric care, AI tools require evidence of clinical validity, transportability, workflow compatibility, patient acceptability, equity, and appropriate governance rather than technical performance alone.

Highlights

  • Psychiatric AI has progressed through successive technical waves, but routine clinical implementation remains limited.

  • Persistent barriers include uncertain clinical targets, limited external validation, poor transportability, workflow integration, equity, patient trust, and governance.

  • Future clinical impact depends on prospective validation, implementation trials, patient‐centered evaluation, and mental health–specific governance.


Artificial intelligence (AI) now occupies a central position in psychiatric research, yet its relationship with the field is considerably older and more complex than the recent surge of interest in machine learning and generative systems suggests. When contemporary researchers apply deep learning to diagnostic classification, use natural language processing to analyze clinical documentation, or consider large language models (LLMs) for administrative or clinical support, they are working within a longer tradition of computational ambition that extends back more than five decades. Recovering this history is not merely an exercise in chronology. It is a way to understand why certain promises recur, why technical advances often lose force when exposed to clinical complexity, and why the distance between algorithmic performance and routine psychiatric practice has remained substantial across multiple technological eras (1, 2).

Psychiatry presents computational science with a domain of unusual difficulty. Unlike many other areas of medicine, it has historically lacked objective biomarkers, stable imaging signatures, laboratory measures, or cellular assays that define most clinical targets with high specificity. Psychiatric diagnoses remain anchored in behavioral observation, clinical interviews, narrative formulation, and categorical classification systems whose boundaries are often porous and whose biological correlates remain incomplete. This combination has made psychiatry especially attractive to computational methods promising pattern detection, stratification, and prediction, while simultaneously making it difficult to translate model performance into robust clinical utility. Across five decades, psychiatric AI has therefore been shaped by a recurring tension: the field seeks computationally tractable representations of mental disorders, but the targets themselves are heterogeneous, context‐dependent, biologically under‐specified, and embedded in relational care (3, 4, 5, 6).

Many prior reviews have addressed important components of this landscape, including AI applications in mental health care, machine learning for diagnosis or monitoring, computational psychiatry, digital phenotyping, precision psychiatry, digital mental health, and LLMs (1, 2, 3, 4, 5, 7, 8, 9, 10, 11). These reviews have clarified contemporary opportunities and limitations, particularly weak external validation, poor transportability, diagnostic heterogeneity, digital phenotyping concerns, and governance gaps. However, many are organized around specific technologies, modalities, disorders, or recent digital mental health applications. Less attention has been given to the cross‐era pattern that links early symbolic simulation, expert systems, connectionist models, supervised learning, computational psychiatry, digital phenotyping, multimodal monitoring, and generative AI.

This review addresses that gap by offering a cross‐era synthesis of psychiatric AI from 1972 to 2025. Its central argument is not that the field has failed to progress, but that progress has been cyclical rather than linear. Each technological wave has expanded computational sophistication while re‐encountering a familiar set of translational constraints: heterogeneous clinical targets, limited external validation, poor transportability, uncertain interpretability, weak workflow integration, inequitable generalizability, and incomplete governance. From PARRY to LLMs, psychiatric AI has repeatedly shown that computational novelty alone is insufficient when the objects of prediction remain clinically complex, context‐dependent, and only partly biologically specified.

This narrative review examines how major AI paradigms emerged across successive historical periods, how they were linked to psychiatric research and clinical ambitions, and why translation into routine care has remained limited despite substantial technical advances. By situating contemporary AI within this longer history, the review aims to clarify where psychiatric AI may contribute, where the evidence remains insufficient, and what forms of validation, implementation, equity assessment, and governance are needed before AI systems can responsibly support mental health care.

METHODOLOGY

Study Design

This article is a narrative historical review of AI in psychiatry. A narrative design was chosen because the objective was to synthesize a heterogeneous interdisciplinary literature spanning symbolic AI, expert systems, connectionist models, supervised machine learning, computational psychiatry, digital phenotyping, multimodal monitoring, and generative AI. The review was designed to examine how successive computational paradigms have been applied to psychiatric research and care, and why recurrent translational barriers have persisted across eras.

The review was structured as an interpretive, historically organized synthesis rather than as a systematic review or meta‐analysis. No protocol was registered, no formal risk‐of‐bias assessment was performed, and no quantitative pooling was conducted. To improve transparency, we specified the information sources, search concepts, source‐selection criteria, data‐charting domains, and synthesis approach.

Information Sources and Search Strategy

Sources were identified through PubMed/MEDLINE, Web of Science Core Collection, Google Scholar, ScienceDirect, backward and forward citation tracking, and targeted searches for historically important or difficult‐to‐index material. PubMed/MEDLINE and Web of Science Core Collection were used as the main bibliographic databases for biomedical, psychiatric, and interdisciplinary literature. Google Scholar was used as a complementary discovery tool to identify citation links, older publications, historical reviews, policy documents, and materials not consistently indexed in biomedical databases. ScienceDirect was used as a supplementary retrieval platform for relevant full‐text articles identified through database searches, citation tracking, or targeted historical searches. This approach was intended to provide structured and transparent source identification, not exhaustive evidence capture, and should not be interpreted as a reproducible systematic or scoping review search strategy.

Searches were conducted primarily in English, and the synthesis relied mainly on English‐language sources. Non‐English literature was not systematically searched, although this limitation was considered when interpreting issues of global generalizability, multilingual deployment, and applicability to non‐Western or resource‐limited settings. Database searches were conducted from database inception through December 31, 2025, to capture early antecedent work as well as contemporary developments. The historical synthesis focused on the period from 1972 to 2025. The final search was completed on April 7, 2026.

Because this was a narrative historical review rather than a systematic review, source identification was iterative and concept‐driven rather than based on a single exhaustive Boolean strategy. Search concepts were selected to reflect changes in terminology across eras, including symbolic AI, expert systems, knowledge‐based systems, connectionist models, machine learning, deep learning, computational psychiatry, digital phenotyping, natural language processing, speech and language biomarkers, and LLMs. These concepts were combined with psychiatry‐ and mental health–related terms, including psychiatry, mental health, psychiatric disorders, schizophrenia, depression, suicide risk, psychosis, and bipolar disorder.

Targeted searches were also conducted for specific historical or translational domains unlikely to be fully captured by broad database searches. Examples included “PARRY artificial paranoia,” “Psyxpert psychiatry expert system,” “DSM‐III expert systems psychiatry,” “support vector machine neuroimaging psychiatry,” “electronic health record suicide risk prediction,” “digital mental health implementation,” “clinical note natural language processing psychiatry,” “social media mental health prediction,” “AI regulation mental health,” and “LLMs psychiatry.”

Source Selection

Sources were selected purposively according to their relevance to the review's historical and translational objectives. Priority was given to sources that introduced or exemplified a major AI paradigm in psychiatry; described a clinically relevant application, such as diagnosis, risk prediction, treatment‐response prediction, digital monitoring, documentation, or decision support; addressed validation, transportability, implementation, interpretability, safety, equity, or governance; synthesized an important subfield through a high‐quality review or meta‐analysis; or clarified historically important developments in psychiatric AI, diagnostic formalization, or computational psychiatry.

Sources were generally excluded when they were purely technical AI papers without a clear psychiatric or mental health application, commentaries without substantial historical or translational relevance, nonmedical AI papers unrelated to mental health, duplicate reports of the same study, or articles whose contribution was already represented by a more comprehensive primary source or review. Because the aim was representative historical synthesis rather than exhaustive evidence enumeration, source selection was interpretive rather than PRISMA‐based, and no formal screening flow diagram or quantitative eligibility count was generated.

Data Charting

For each source considered central to the synthesis, the authors charted the following information when available: historical era, representative study or source, AI paradigm or method, psychiatric task or research objective, data modality, validation approach, implementation status, key translational limitation, and relevance to the review's central argument. These categories informed both the narrative synthesis and the structured evidence table. For Table 1, 36 representative sources were centrally charted to summarize cross‐era milestones, AI paradigms, clinical or research tasks, data modalities, validation approaches, implementation status, and recurring translational limitations. Source selection and data charting were performed by the corresponding author and reviewed by the coauthors, with interpretive disagreements or uncertainties resolved by discussion among the authors.

TABLE 1.

Representative milestones in psychiatric AI and recurring translational barriers, 1972–2025. a

Era Representative study/source AI paradigm or method Clinical/research task Data modality Validation approach Implementation status Key translational limitation
1970s Colby et al., “Artificial paranoia”; Colby et al., “Turing‐like validation of PARRY” (12, 13) Symbolic simulation Simulation of paranoid communication and psychiatric interview behavior Typed interview dialog Turing‐like indistinguishability testing with psychiatrists Laboratory demonstration; not implemented in routine care Conversational plausibility did not establish diagnostic validity, clinical utility, or treatment relevance
1980s Overby, “Psyxpert: expert‐system prototype” (14) Rule‐based expert system; production rules; backward chaining Diagnostic support for psychotic presentations Structured clinical inputs Prototype‐level evaluation Experimental decision‐support prototype; no routine deployment Narrow scope, rule brittleness, and limited applicability to ambiguous or heterogeneous clinical presentations
1990s Clarke et al., “Monash Interview for Liaison Psychiatry”; Cohen and Servan‐Schreiber; Braver et al.; Hoffman and McGlashan (15, 16, 17, 18, 19) Computer‐assisted diagnostic algorithm; connectionist models Structured diagnosis; mechanistic modeling of cognition, language, and psychosis Interview‐based clinical data; simulated neural‐network representations Reliability/procedural validity for structured interview; conceptual and computational plausibility for models Research and structured‐assessment use; not autonomous clinical AI Constrained clinical inputs and mechanistic abstraction limited bedside translation
2000s Davatzikos et al.; Orrù et al.; Kambeitz et al.; Nielsen et al. (20, 21, 22, 23) Supervised machine learning; support vector machines; multivariate pattern recognition Individual‐level diagnostic classification in schizophrenia and depression Structural MRI, functional MRI, diffusion imaging Mostly internal validation and cross‐validation; variable external validation Research‐stage biomarker development; not routine clinical use Small samples, site effects, preprocessing heterogeneity, overfitting risk, and limited transportability
2010–2015 Montague et al.; Wang and Krystal; Friston et al.; Huys et al. (7, 24, 25, 26) Computational psychiatry; Bayesian inference; predictive coding; reinforcement learning Mechanistic modeling of symptoms, learning, valuation, and belief updating Behavioral, cognitive, neurobiological, and computational model parameters Conceptual and model‐based validation; limited direct clinical testing Research framework; limited clinical deployment Mechanistic sophistication remained difficult to translate into actionable clinical tools
2010–2015 Tran et al.; van Dinteren et al.; Widge et al. (27, 28, 29) Machine learning using electronic medical records; electrophysiological biomarkers Suicide‐risk prediction; antidepressant treatment‐response prediction Structured EHR data; event‐related potentials/electroencephalography Retrospective evaluation; biomarker prediction studies and later meta‐analytic evaluation Proof‐of‐concept prediction; not routine care Retrospective design, unstable outcome definitions, limited out‐of‐sample replication, and uncertain clinical actionability
2015–2020 Chekroud et al.; Barak‐Corren et al.; Salazar De Pablo et al. (30, 31, 32) Machine learning for precision psychiatry and large‐scale risk prediction Antidepressant remission prediction; incident suicide‐attempt prediction; individualized prediction models Baseline clinical variables; structured EHR data; clinical datasets Cross‐trial and multisite validation in selected studies; broader literature showed limited external validation Research‐stage and health‐system prediction; implementation remained uncommon Modest performance, calibration concerns, false‐positive burden, limited implementation, and uncertain effect on clinical outcomes
2015–2020 Onnela and Rauch; Huckvale et al. (33, 34) Digital phenotyping; smartphone‐based behavioral sensing Continuous monitoring of behavior, mobility, sleep, activity, and social functioning Smartphones, wearables, passive sensing, ecological momentary assessment Feasibility and observational studies; heterogeneous validation Rapidly expanding research field; limited routine clinical adoption Data quality, missingness, privacy, safety, bias, and unclear clinical purpose
2020–2025 Garriga et al.; Guerreiro et al.; Bufano et al.; Khoo et al. (35, 36, 37, 38) Multimodal machine learning; passive sensing; longitudinal EHR modeling Prediction of mental health crises and dynamic clinical states EHR, smartphone, wearable, behavioral, and multimodal data streams Longitudinal validation; selected cross‐site or transatlantic transferability studies Emerging implementation research; not broadly standardized Transportability, calibration across systems, workflow integration, and sustainability
2020–2025 Torous et al.; Hua et al.; Jin et al.; Wang et al.; Wang et al.; Artsi et al.; Martinez‐Martin et al.; Shumate et al.; FDA guidance (9, 10, 11, 39, 40, 41, 42, 43) Digital mental health, large language models, generative AI, and AI‐enabled software governance Screening support, documentation, counseling support, clinical assistance, digital interventions, and oversight Natural language, clinical documentation, conversational inputs, digital platforms, policy and regulatory documents Early scoping/systematic reviews; heterogeneous evaluation standards; policy and regulatory analysis Exploratory or early implementation; governance remains emergent and fragmented Hallucinations, bias, privacy, safety, accountability, limited mental‐health‐specific evaluation standards, and uncertain regulatory pathways
a

AI, artificial intelligence; EHR, electronic health record; FDA, U.S. Food and Drug Administration; LLM, large language model; MRI, magnetic resonance imaging.

The charting process emphasized translational relevance rather than algorithmic detail alone. Studies were interpreted in relation to whether they addressed, reproduced, or failed to resolve persistent barriers such as target validity, diagnostic heterogeneity, external validation, transportability, interpretability, workflow integration, equity, safety, and governance.

Narrative Synthesis

The literature was synthesized qualitatively using a chronological and translational framework. The field was organized into broad eras: early symbolic simulation and PARRY in the 1970s; expert systems and nosological formalization in the 1980s; connectionist models and computer‐assisted structured diagnosis in the 1990s; supervised machine learning and neuroimaging prediction in the 2000s; computational psychiatry and early clinical prediction models from 2010 to 2015; digital phenotyping, precision psychiatry, and external‐validation concerns from 2015 to 2020; and multimodal monitoring, generative AI, and governance from 2020 to 2025.

Within each era, the synthesis focused on three questions: what computational advance became prominent, what psychiatric task or clinical ambition it was used to address, and what translational limitations persisted despite technical progress. This approach was intended to clarify the review's central thesis: psychiatric AI has repeatedly expanded computational sophistication while recurring barriers related to target validity, validation, transportability, implementation, equity, and governance have remained incompletely resolved.

Historical and Governance Sources

Targeted historical searches were used because early psychiatric AI, including symbolic simulation and expert‐system prototypes, is not consistently captured by contemporary biomedical indexing. Historical sources were included when they were directly relevant to the development of psychiatric AI or to the formalization of psychiatric diagnosis and could be verified through journal archives, publisher websites, PubMed records, institutional sources, or authoritative secondary literature.

Governance and regulatory sources were included when they addressed AI‐enabled clinical software, digital mental health, LLMs, privacy, accountability, safety, or mental health–specific oversight. These sources were used to contextualize contemporary translational challenges rather than to provide a comprehensive legal or regulatory review.

Methodological Limitations

This review has limitations inherent to narrative historical synthesis. Although the search strategy was structured and transparent, it was not designed to identify every eligible publication, generate a PRISMA‐style screening flow, or support quantitative pooling. Source selection was purposive and interpretive, which may introduce selection bias despite the use of explicit selection criteria. The long period covered includes major changes in psychiatric nosology, AI terminology, data modalities, validation standards, and regulatory expectations, limiting direct comparability across eras. The review also emphasized psychiatric, biomedical, translational, and governance literature; therefore, some purely technical computer‐science contributions may be underrepresented.

In addition, because searches were conducted primarily in English, relevant non‐English literature may be underrepresented. This is particularly important for domains involving multilingual natural language processing, speech and language biomarkers, digital mental health implementation, and the generalizability of psychiatric AI in non‐Western or resource‐limited settings. Finally, because AI in psychiatry is evolving rapidly, especially in relation to generative AI, digital mental health, and regulatory oversight, conclusions regarding clinical readiness and governance may require updating as new evidence emerges.

NARRATIVE SYNTHESIS

Table 1 provides a structured overview of representative milestones in psychiatric AI from 1972 to 2025. Rather than serving as an exhaustive inventory, the table is intended to orient the reader to the major computational paradigms, clinical or research tasks, data modalities, validation approaches, implementation status, and recurring translational limitations that organize the narrative synthesis.

Early Simulation and Symbolic Modeling: PARRY (1972–1980)

The earliest phase of AI in psychiatry was driven less by immediate clinical deployment than by a theoretical problem: whether symbolic computation could model the observable communicative features of mental illness. In Artificial Paranoia (1971), Colby, Weber, and Hilf presented “artificial paranoia” as a computer simulation model and a theoretical account of paranoid communicative behavior, explicitly proposing indistinguishability from human interviews as the criterion of success. In the 1972 follow‐up, Colby and colleagues framed the project within the computer simulation of human mental functions and treated validation as a methodological problem of model justification.

PARRY, developed by Kenneth Mark Colby and colleagues at Stanford, was the central artifact of this period. In Artificial Paranoia (1971), it was presented as a computer simulation of paranoid processes designed for use in a diagnostic psychiatric interview. Rather than operating as mere surface‐level text mimicry, the model was organized around an internal representational structure: a delusional belief system that governed conversational strategy, together with affective‐state variables (fear, anger, and mistrust) that changed in response to perceived malevolence and shaped whether the system counterattacked, withdrew, or produced noncommittal replies. Colby and coauthors explicitly framed it as a theoretical model intended to explain paranoid communicative behavior through its inner structure, not just imitate it externally (12, 13).

In the subsequent validation study, Turing‐Like Indistinguishability Tests for the Validation of a Computer Simulation of Paranoid Processes (1972), Colby and colleagues evaluated whether psychiatrists could distinguish teletyped interviews with the program from interviews with real patients. The most precise way to report the result is not to say “48% accuracy,” but rather that psychiatrists performed at approximately chance level: 6 of 8 interview judges made the wrong identification, and in a follow‐up protocol test the responses yielded 21 correct identifications and 19 incorrect ones out of 40. Accordingly, the result can be stated either as 52.5% correct identification or as a 47.5% misidentification rate, with the key point being that trained psychiatrists did not reliably distinguish the simulation from real patients under those test conditions (12, 13).

However, the legacy of the 1970s remains conceptual rather than instrumental. The systems were isolated laboratory demonstrations, constrained by a lack of shared datasets and a narrow focus on symbolic representation. While these models lacked a realistic pathway to clinical decision support or treatment, they established a premise that endures today: that psychiatric phenomena, however complex, are inherently tractable to formal representation as information processing. Rather than representing a direct precursor of contemporary clinical AI, these early systems are better understood as conceptually important attempts to formalize aspects of psychiatric phenomena within computational frameworks. In that limited but significant sense, they anticipated later efforts to model, classify, and predict mental health–related states using formal representation.

1980–1990: The Formalization of Nosology and the Expert System Paradigm

During the 1980s, the central development in psychiatry was not AI, but the consolidation of a more standardized diagnostic framework, especially after the publication of DSM‐III in 1980. Historical analyses describe DSM‐III as a major turning point that moved American psychiatry toward explicit diagnostic criteria, symptom‐based descriptions, and greater reliability across clinicians. In that sense, the decade was primarily devoted to establishing the conceptual and diagnostic foundations that would later make more formal computational approaches conceivable (44, 45).

Within this context, AI did not yet play a leading role in routine psychiatric practice. Its presence was limited and largely experimental, usually in the form of expert‐system prototypes rather than broadly implemented clinical tools. One of the clearest examples was Psyxpert (1987), an expert system prototype designed to assist psychiatrists in diagnosing mental disorders when psychotic features predominated in the clinical presentation. According to the original report, its knowledge base was represented as production rules, and the system used a backward‐chaining control strategy, a menu‐driven interface, and an explanation subsystem to support the consultation process (14).

Methodologically, these early systems were important less for their clinical impact than for what they revealed about the limits of symbolic AI in psychiatry. Although systems such as Psyxpert showed that parts of diagnostic reasoning could be formalized into explicit rules, their scope remained narrow and their usefulness depended on relatively constrained clinical scenarios. More broadly, the expert‐system paradigm of the 1980s was powerful for well‐defined domains, but it was also vulnerable to brittleness when confronted with complex, ambiguous, or heterogeneous real‐world cases, a limitation widely recognized in the history of AI in medicine (46, 47).

1990–2000: The Neurobiological Turn and the Connectionist Paradigm

During the 1990s, AI did not become a dominant clinical force in psychiatry. Instead, the decade was shaped more strongly by the growing neurobiological and cognitive‐neuroscience orientation of psychiatric research especially in schizophrenia, while computational work developed along two more limited paths: computer‐assisted structured diagnosis in clinical settings and connectionist models used to explore mechanisms of cognition and psychosis. In this sense, the period did not consolidate AI as a routine psychiatric technology, but it did broaden the field beyond symbolic expert systems by introducing algorithm‐assisted interviewing and neural‐network‐based hypothesis testing (15, 16, 24).

On the clinical side, the Monash Interview for Liaison Psychiatry (MILP), reported in 1998, was a structured interview for patients with physical and psychiatric comorbidity that was linked to a computerized diagnostic algorithm capable of generating DSM‐III‐R, ICD‐10, and DSM‐IV diagnoses. On the research side, the most visible AI‐related work came from connectionist and parallel distributed processing models of schizophrenia. Cohen and Servan‐Schreiber used connectionist models to relate prefrontal and mesocortical dopamine dysfunction to deficits in attention and language processing, Braver, Barch, and Cohen later extended this framework to cognitive control, and Hoffman and colleagues used neural‐network simulations to examine how altered cortical connectivity or excessive pruning could generate psychosis‐like phenomena, including hallucination‐like outputs and delusion‐like recurrent representations (15, 16, 17, 18, 19).

These developments nevertheless had clear methodological limits: MILP was better understood as an algorithm‐assisted structured interview than as autonomous AI, and its usefulness depended on constrained clinical inputs and a specific consultation‐liaison setting. The connectionist studies were even further from bedside deployment: they were primarily heuristic research models, often centered on schizophrenia, designed to test mechanistic hypotheses rather than to provide broad diagnostic or treatment support. Retrospective reviews of computational psychiatry and schizophrenia modeling suggest that these approaches generated plausible abstractions and multiple competing hypotheses, but remained narrow in scope and only partially translatable to routine practice (15, 24, 48).

2000–2010: The Predictive Turn and the Rise of Supervised Machine Learning

During the 2000s, AI in psychiatry entered a more explicitly predictive phase, although this shift was gradual rather than abrupt. The central development was not the routine clinical use of AI, but the growing application of supervised machine‐learning methods to high‐dimensional neuroimaging data, especially structural Magnetic Resonance Imaging (MRI) and functional MRI (fMRI). In contrast to the symbolic expert systems of the 1980s and the mechanistic connectionist models of the 1990s, the dominant ambition of this period was to classify individual patients from multivariate biological patterns. Later reviews of the field describe this movement as a response to the limited clinical translation of traditional group‐level neuroimaging findings and identify support vector machines (SVMs) as the most visible methodological framework for this new line of work (22, 23).

The clearest early expression of this turn appeared in schizophrenia research. A landmark study by Davatzikos and colleagues in 2005 showed that whole‐brain morphometric MRI analysis could identify a spatially distributed pattern of abnormalities in schizophrenia and that pattern classification achieved high sensitivity and specificity, suggesting possible utility as a diagnostic aid. Subsequent meta‐analytic work confirmed that this was not an isolated result: across 38 schizophrenia studies, neuroimaging‐based multivariate classifiers achieved an overall sensitivity and specificity of about 80%, with resting‐state fMRI showing somewhat higher sensitivity than structural MRI. In historical terms, the major contribution of the decade was therefore not the deployment of mature psychiatric AI in the clinic, but the demonstration that psychiatric neuroimaging could be reframed as an individual‐level prediction problem (20, 21, 23, 49).

Beyond schizophrenia, the same predictive logic began to spread to other psychiatric disorders, although more unevenly and usually later in the decade. Retrospective meta‐analysis in major depressive disorder identified 33 neuroimaging samples and found overall diagnostic performance of 77% sensitivity and 78% specificity, with resting‐state MRI and diffusion tensor imaging generally outperforming structural MRI and task‐based fMRI. These studies also show that SVM was the most common algorithm, while other methods such as Gaussian‐process classifiers, neural networks, random forests, and decision trees appeared only sporadically. Accordingly, the decade is best characterized not as the moment when psychiatry fully embraced clinical AI, but as the period in which supervised learning became a credible research strategy for seeking neurobiological signatures of mental disorders at the level of the individual patient (21, 22, 23, 49).

Methodologically, however, this literature remained fragile. Later syntheses consistently emphasize small or modest samples, heterogeneous preprocessing pipelines, inconsistent cross‐validation schemes, and the risk of overfitting or information leakage when feature selection was not properly nested within model validation. Clinical heterogeneity posed an additional problem: age, medication status, illness stage, symptom profile, and diagnostic unreliability all had the potential to influence classification performance, while true cross‐site generalizability was rarely demonstrated. For that reason, even when accuracy looked promising in development cohorts, most models were not yet ready for routine clinical adoption. What the 2000s established, then, was less a solved diagnostic technology than a durable research agenda: psychiatry could be studied through prediction, but predictive performance alone did not guarantee translational usefulness (21, 23, 49).

2010–2015: The Formalization of Computational Psychiatry and the Mechanistic‐Predictive Hybrid

Between 2010 and 2015, psychiatry did not simply adopt “more AI”; rather, this was the period in which computational psychiatry became clearly recognized as a named interdisciplinary field. Seminal reviews from this interval described it as a new framework for characterizing mental dysfunction in terms of aberrant computations, while later syntheses clarified that the field was developing along two complementary lines: theory‐driven models aimed at explaining mechanisms and data‐driven models aimed at improving prediction. For that reason, the period is best understood not as a sudden methodological revolution, but as a consolidation phase in which mechanistic and predictive approaches began to coexist within a shared conceptual vocabulary (7, 24, 25).

On the mechanistic side, the most influential developments came from Bayesian inference, predictive coding, reinforcement learning, and active inference. Reviews by Montague, Friston, and colleagues argued that psychiatric symptoms could be studied as disturbances in belief updating, valuation, and learning, while Friston and Stephan explicitly framed computational psychiatry as a way to reinterpret symptoms through disordered inference and maladaptive prediction‐error processing. At the same time, reinforcement‐learning approaches associated with Dayan and others linked abnormalities in reward prediction and dopaminergic signaling to clinically relevant phenomena such as addiction, depression, and impaired decision‐making. These models did not yet produce routine clinical tools, but they gave the field a more rigorous explanatory language than earlier symbolic systems or loosely specified neurobiological metaphors (7, 24, 25, 26).

In parallel, the predictive agenda became more clinically oriented. As Electronic Health Records (EHR) and larger clinical datasets became more available, machine‐learning models were increasingly applied to outcomes that mattered directly for care rather than only to case‐control classification. A notable example was the 2014 study showing that models built from electronic medical records could predict short‐term suicide risk better than clinician assessment alone. At the same time, biomarker‐oriented work on treatment response also gained visibility: studies such as the 2015 iSPOT‐D report examined whether electrophysiological markers could help predict antidepressant response, reflecting the broader ambition to move psychiatry toward more individualized forecasting. Even so, these applications were still emerging rather than routine, and their value was primarily proof‐of‐concept (27, 28).

Methodologically, however, the translational gap remained substantial. Reviews from and after this period repeatedly emphasized that mechanistic models were elegant but difficult to validate directly at the bedside, whereas predictive models were often retrospective, dependent on unstable outcome definitions, and vulnerable to overfitting, site effects, and limited external generalizability. The problem was especially visible in biomarker studies for treatment selection, where later evaluations concluded that apparently promising signals often lacked robust out‐of‐sample replication. Thus, by 2015, psychiatric AI had become more conceptually sophisticated and clinically ambitious, but it was still better described as a dual‐track research program, one explanatory and one predictive, than as a mature set of tools ready for routine clinical implementation. These limitations were reinforced by persistent challenges in external validity, outcome definition, and generalizability across settings (26, 29, 50).

2015–2020: Digital Phenotyping, Precision Psychiatry, and the Push for Real‐World Validity

Between 2015 and 2020, AI in psychiatry expanded substantially in both scope and ambition, but this was less a clean methodological break than a broadening of the field. Classical supervised machine‐learning approaches remained central, especially for diagnosis, prognosis, and treatment prediction, while deep‐learning methods began to supplement them in areas such as neuroimaging analysis. At the same time, precision psychiatry became an explicit aspirational framework, framed as a move toward individualized prediction and stratified care rather than purely group‐level inference (8, 51).

A defining development of this period was the rise of digital phenotyping. In 2016, Onnela and Rauch defined it as the “moment‐by‐moment” in situ quantification of the human phenotype using data from personal digital devices, and by 2019 reviews were already describing more than 80 peer‐reviewed psychiatric digital‐phenotyping publications since 2015. Smartphones, wearables, and related tools made it possible to capture behavior, sleep, activity, mobility, and social functioning outside the clinic, shifting psychiatric assessment from episodic observation toward more continuous monitoring. At the same time, this literature quickly raised practical questions about quality, safety, bias, and clinical feasibility (33, 34).

In parallel, the predictive agenda became more clinically consequential. Rather than focusing only on case‐control classification, studies increasingly aimed to forecast outcomes such as antidepressant remission, relapse, and suicidal behavior. A notable example was Chekroud and colleagues' cross‐trial machine‐learning model for predicting remission after antidepressant treatment: using 25 variables selected from 164 patient‐reportable baseline measures, the model predicted remission in STAR*D with 64.6% accuracy and retained significant, above‐chance performance in external validation in the escitalopram arm of the independent COMED trial (59.6% accuracy; p = 0.043). Meanwhile, EHR‐based suicide‐risk models achieved large‐scale multisite validation: a 2020 study across five US health systems reported prediction of incident suicide attempts using structured EHR data from more than 3.7 million patients. Toward the end of the decade, normative modeling also emerged more clearly as an alternative to simple case‐control contrasts, emphasizing individualized deviation from expected biological variation rather than average group differences alone (30, 31, 52).

However, the expansion of the field was matched by stronger methodological scrutiny. Reviews from this period emphasized that many psychiatric prediction studies still relied on small or selected samples, weak external validation, heterogeneous outcomes, and substantial risks of overfitting, data leakage, and poor transportability. A 2020 systematic review of precision psychiatry found 584 prediction‐modeling studies, but only 10.4% included internal validation, 4.6% external validation, and just 0.2% had progressed to implementation. By 2020, the main question was no longer simply whether psychiatric AI models could be built, but whether they could meet the evidentiary standards required to be trustworthy, generalizable, and clinically useful (32, 34).

2020–2025: Multimodal Monitoring, Generative AI, and the Turn to Governance

Between 2020 and 2025, AI in psychiatry underwent rapid expansion in both methodological breadth and clinical ambition, although this growth still outpaced routine implementation. Rather than marking a completed transition to clinical maturity, the period is better understood as one of intensification: precision psychiatry became a more explicit organizing aspiration, digital mental health broadened beyond earlier app‐based models, and multimodal as well as generative systems moved closer to real clinical workflows (2, 9).

A major development of this period was the scaling of multimodal monitoring outside the clinic. Digital phenotyping and passive sensing approaches increasingly combined smartphone‐derived signals, wearable data, ecological momentary assessments, and other behavioral streams to capture mood, sleep, mobility, and social functioning in more continuous and naturalistic settings. Systematic reviews from this period describe a rapid growth of multimodal sensing studies, while longitudinal EHR‐based models demonstrated that machine learning could continuously monitor patients and predict short‐term mental health crises, reinforcing a view of psychiatric states as dynamic rather than static (35, 36, 37, 38).

A second defining development was the emergence of generative AI, especially LLMs, from late 2022 and particularly from 2023 onward. In a field heavily dependent on language, narrative, and clinical documentation, LLMs were explored for mental health screening, counseling support, documentation, and the analysis of unstructured psychiatric text. Reviews published in 2025 suggest substantial promise, but they also emphasize that the evidence base remains early, evaluation standards are still heterogeneous, and real‐world implementation in psychiatry is limited (10, 11, 40, 41).

At the same time, the central challenge of the field shifted more clearly toward governance. Concerns about hallucinations, bias, privacy, consent, accountability, and clinical safety became more prominent as models grew more powerful and more conversational. Ethical analyses of digital phenotyping had already identified privacy, bias, and accountability as foundational concerns, and by 2025 reviews of LLMs in mental health and legislative analyses in the United States were describing a fragmented oversight landscape, with mental health–specific regulation still limited. Thus, the significance of 2020–2025 lies not only in technical acceleration, but in the fact that psychiatric AI entered a phase where questions of validity, safety, and governance became as important as performance itself (9, 42, 43, 53).

Domains at the Translational Frontier: NLP, Speech, Digital Therapeutics, Social Media, and Equity

Several rapidly developing domains illustrate both the promise and the unresolved translational problems of contemporary psychiatric AI. Clinical‐note natural language processing is especially relevant in psychiatry because much of the phenotype is encoded in narrative documentation rather than laboratory values or imaging biomarkers. NLP systems can extract symptoms, functioning, risk‐related information, treatment trajectories, and contextual features from unstructured mental health records, but they also inherit documentation bias, institutional language patterns, missing context, privacy risks, and limited transportability across health systems (54, 55). Speech and language biomarkers are similarly aligned with psychiatric assessment, given the clinical importance of prosody, fluency, semantic coherence, affective expression, and interpersonal communication. However, their performance may vary by language, culture, recording conditions, demographic factors, and disorder heterogeneity, making multilingual, cross‐cultural, and longitudinal validation essential before clinical use (56, 57).

Social‐media‐based prediction represents a different frontier, where behavioral and linguistic signals may be available at large scale but are often detached from clinical assessment. These approaches raise unresolved questions about construct validity, consent, representativeness, platform drift, false‐positive harms, crisis‐response responsibility, privacy, and the boundary between care and surveillance (58, 59). Digital therapeutics and AI‐supported digital mental health tools extend the field from prediction toward intervention, including symptom monitoring, self‐management, engagement support, asynchronous care, chatbot‐supported interactions, and potentially adaptive treatment delivery. However, their relevance for routine psychiatric care depends on more than usability or algorithmic personalization; it requires prospective evaluation, safety monitoring, clinically meaningful outcomes, reimbursement pathways, and integration with existing care models (60, 61).

These domains also make implementation, patient perspectives, and equity more central. Psychiatric AI systems should be evaluated not only for discrimination, calibration, or benchmark performance, but also for workflow fit, clinician response, patient trust, therapeutic relationship effects, and unintended harms. Reporting frameworks for AI trials and early‐stage clinical evaluation provide useful standards for assessing AI‐enabled interventions beyond retrospective performance metrics (62, 63). Patient perspectives are also essential because acceptability may depend on transparency, confidentiality, perceived accuracy, human oversight, and whether AI is viewed as supporting rather than replacing clinical care (64).

This is particularly important in non‐Western, multilingual, and resource‐limited settings, where models trained in high‐income, English‐language, highly digitized health systems may not generalize to different documentation practices, cultural expressions of distress, service pathways, or infrastructure constraints (65). Regulatory pathways for AI‐enabled software and generative systems are still developing, and LLMs may require oversight approaches that differ from earlier AI‐enabled medical devices because of their scale, adaptability, and open‐ended outputs (66, 67). For these reasons, the translational frontier of psychiatric AI is not defined by computational capability alone, but by whether systems can be validated, implemented, governed, and trusted across diverse clinical and social contexts. These recurrent translational cycles across successive eras of psychiatric AI are summarized in Figure 1.

FIGURE 1.

FIGURE 1

Recurrent translational cycles in psychiatric AI, 1972–2025. Panel A summarizes successive technological waves in psychiatric AI from symbolic simulation to multimodal and generative systems. Panel B reframes this history as a recurrent cycle in which computational advances generate proof‐of‐concept performance, encounter validation and implementation friction, and are followed by renewed paradigms or data modalities. The figure emphasizes that technical sophistication has increased across eras, while core translational barriers—including target validity, diagnostic heterogeneity, external validation, transportability, interpretability, workflow integration, equity, safety, and governance—have persisted. AI, artificial intelligence.

DISCUSSION

Principal Findings and Contribution

This narrative review provides a cross‐era translational synthesis of psychiatric AI from 1972 to 2025. Psychiatric AI has moved through recurrent cycles in which new computational paradigms expanded technical capability while leaving core translational barriers incompletely resolved. This review adds a longer historical interpretation: problems such as weak external validation, poor transportability, diagnostic heterogeneity, interpretability, workflow integration, equity, and governance are not isolated limitations of the current machine‐learning or generative‐AI era. They have recurred whenever computational methods have been applied to heterogeneous psychiatric targets. This historical pattern does not imply lack of progress. The field has moved from symbolic simulation and rule‐based diagnostic prototypes to connectionist models, neuroimaging classifiers, computational psychiatry, EHR‐based prediction, digital phenotyping, multimodal monitoring, clinical‐note NLP, speech and language biomarkers, digital therapeutics, and LLMs. These developments have expanded what can be modeled, measured, and predicted. Yet the central translational question has remained largely unchanged: whether apparent computational performance can be converted into clinically useful, externally valid, interpretable, equitable, and governable tools for mental health care.

Why the Translational Gap Persists

The translational gap persists in part because psychiatric AI inherits the unresolved problem of psychiatric target validity. Most psychiatric disorders are not defined by pathognomonic biomarkers, single laboratory values, or discrete lesions. They are syndromic constructs based on symptoms, behavior, clinical narrative, impairment, and longitudinal judgment. Diagnostic categories may be useful for communication and care, but they remain heterogeneous, overlapping, and only incompletely aligned with biological mechanisms (4, 51). As a result, AI systems trained to reproduce diagnostic labels, symptom ratings, or administrative outcomes may learn patterns embedded in clinical practice rather than stable disease mechanisms.

This limitation has appeared in different forms across eras. Expert systems formalized diagnostic rules but were brittle in complex presentations. Neuroimaging classifiers reframed psychiatric diagnosis as an individual‐level prediction problem but were limited by small samples, site effects, preprocessing variation, and inconsistent external validation (20, 21, 22, 23, 48). Computational psychiatry offered a more mechanistic language for learning, valuation, belief updating, and prediction error, but translating elegant models into actionable clinical tools has remained difficult (7, 24, 25, 26). EHR‐based models and precision psychiatry approaches have moved closer to clinically relevant outcomes, including suicide risk and treatment response, yet retrospective performance, calibration, transportability, and clinical utility remain major concerns (7, 27, 28, 29, 30, 31, 32).

A second reason is that the field has often emphasized proof‐of‐concept discrimination before demonstrating clinical value. Internal validation, cross‐validation, benchmark performance, and retrospective prediction are important early steps, but they are not sufficient for implementation. In psychiatry, a prediction may alter clinician behavior, affect patient autonomy, intensify monitoring, influence access to care, or trigger safety interventions. For this reason, psychiatric AI requires not only statistical performance but also prospective evaluation, decision‐impact studies, evidence of benefit, and careful assessment of unintended harms.

From Model Development to Implementation

A central implication of this review is that psychiatric AI should now be judged less by whether models can be built than by whether they can be responsibly implemented. Implementation requires attention to workflow fit, clinician burden, alert fatigue, documentation practices, model updating, accountability, and the availability of effective response pathways. A suicide‐risk model, for example, is not clinically meaningful simply because it stratifies risk. Its usefulness depends on whether clinicians can act on the output, whether interventions are available, whether false positives and false negatives are managed ethically, and whether the model remains calibrated across settings and time.

This standard is particularly important because few individualized prediction models in psychiatry have progressed to robust external validation or implementation‐oriented testing (32, 34). Reporting frameworks for AI clinical trials and early‐stage evaluation of AI decision‐support systems provide useful benchmarks for moving beyond retrospective performance metrics (62, 63). Future studies should therefore measure not only discrimination and calibration, but also clinician response, patient outcomes, workflow integration, safety, equity effects, and sustainability. This shift is essential if psychiatric AI is to move from technical demonstration to clinically meaningful support.

Equity, Patient Trust, and Global Generalizability

The newer domains added to this revision reinforce the same translational pattern. Clinical‐note NLP, speech and language biomarkers, social‐media prediction, and digital therapeutics extend the field beyond structured variables and imaging data, but they do not escape the need for rigorous validation. NLP systems may extract symptoms, functioning, risk signals, and treatment trajectories from unstructured records, yet they remain vulnerable to documentation bias, missing context, privacy risks, and limited transportability across health systems (54, 55). Speech and language biomarkers are closely aligned with psychiatric assessment, but their performance may vary by language, culture, recording environment, demographic factors, and disorder heterogeneity (56, 57). Social‐media approaches provide scale, but raise unresolved questions about construct validity, consent, representativeness, platform drift, privacy, and the boundary between care and surveillance (58, 59). Digital therapeutics and AI‐supported digital mental health tools shift the field from prediction toward intervention, but require evidence of clinical effectiveness, safety, engagement, reimbursement feasibility, and integration with care pathways (60, 61).

Equity and trust are therefore not secondary considerations; they are conditions of translation. Many psychiatric AI systems are developed in high‐income, English‐language, highly digitized health systems. Their outputs may not generalize to non‐Western, multilingual, or resource‐limited settings where documentation practices, service pathways, cultural expressions of distress, and infrastructure differ substantially. Models trained on narrow or unrepresentative data may also reproduce or amplify existing disparities (65). Patient perspectives are equally important. Acceptability may depend on transparency, confidentiality, perceived accuracy, human oversight, and whether AI is experienced as supporting rather than replacing clinical care (64). In psychiatry, where trust, narrative, and therapeutic alliance are central to care, systems that are technically plausible but socially illegible may fail to translate.

Governance in the Generative‐AI Era

Governance has become more urgent as psychiatric AI has moved from static prediction models toward digital phenotyping, multimodal monitoring, and generative systems. Digital phenotyping raises concerns about consent, privacy, data ownership, behavioral surveillance, and secondary uses of sensitive information (43). LLMs intensify these concerns because they operate directly through language, the core medium of psychiatric assessment, documentation, psychotherapy, and clinical formulation. Their risks include hallucinated content, biased responses, inappropriate reassurance or escalation, privacy breaches, and unclear accountability. Current reviews suggest promise for documentation, screening support, counseling support, and workflow assistance, but also emphasize heterogeneous evaluation standards and limited real‐world evidence in mental health care (9, 10, 11, 40, 41, 53).

Regulatory pathways for AI‐enabled software are evolving, but mental health–specific oversight remains fragmented (39, 42, 66, 67). This is consequential because psychiatric AI systems may influence diagnosis, risk assessment, triage, documentation, treatment planning, and patient communication. Governance should therefore extend beyond premarket evaluation to include post‐deployment monitoring, model updating, transparency, auditability, adverse‐event reporting, human oversight, and clear allocation of responsibility. For generative AI, additional safeguards are needed around clinical prompting, documentation accuracy, crisis handling, and the boundary between administrative assistance, clinical decision support, and direct patient interaction.

Limitations

This review has limitations. As a narrative historical review, it was designed to provide an interpretive synthesis rather than an exhaustive or quantitatively pooled assessment of the psychiatric AI literature. Although the revised Methods provide a more transparent account of information sources, search strategy, source selection, data charting, and synthesis approach, source selection remained purposive and may have introduced selection bias. The long historical period covered, from 1972 to 2025, includes substantial changes in psychiatric nosology, AI terminology, data modalities, validation standards, and regulatory expectations, which limits direct comparability across eras.

The review also emphasized psychiatric, biomedical, translational, and governance literature rather than the full technical computer‐science literature on AI. This focus is consistent with the clinical objective of the review, but some technical developments may be underrepresented. In addition, rapidly developing domains such as clinical‐note NLP, speech and language biomarkers, social‐media prediction, digital therapeutics, and LLM‐based clinical support require more detailed domain‐specific reviews than could be provided here. Finally, because psychiatric AI continues to evolve rapidly, conclusions regarding clinical readiness, governance, and implementation may require updating as new evidence emerges.

CONCLUSION

Across five decades, psychiatric AI has produced substantial technical advances but limited routine clinical translation. The central lesson is not that AI lacks relevance for psychiatry, but that computational innovation alone is insufficient. Clinical impact will likely depend on clearer targets, robust external and prospective validation, implementation trials, patient‐centered evaluation, equity‐sensitive generalizability, and governance frameworks tailored to the ethical and relational realities of mental health care. Psychiatric AI may become clinically useful, but only when technical capability is matched by evidence of safety, utility, trustworthiness, and responsible implementation.

Zavaleta‐Monestel E, Herrera‐Jiménez LG, Suárez‐Sánchez S, Arguedas‐Chacón S, Mora‐Jiménez J, Millán‐González R. Artificial Intelligence in Psychiatry: Five Decades of Progress and Persistent Translational Challenges. Psych Res Clin Pract. 2026;1–14. 10.1176/appi.prcp.20260059

This work has not been previously presented at any meeting.

This work was conducted at the Health Research Department, Clínica Bíblica, San José, Costa Rica.

Drs. Zavaleta‐Monestel, Herrera‐Jiménez, Suárez‐Sánchez, Arguedas‐Chacón, Mora‐Jiménez, and Millán‐González report no financial relationships with commercial interests.

Artificial intelligence–assisted tools were used during manuscript revision to support language editing, copyediting, and organization of responses to reviewer comments. These tools were not used to conduct source screening, determine source eligibility, extract or chart data independently, generate references, or replace author judgment. All AI‐assisted outputs were critically reviewed, edited, and verified by the authors, who take full responsibility for the accuracy, integrity, and final content of the manuscript.

REFERENCES

  • 1. Ali M, Ali S, Abbas Q, Abbas Z, Lee SW. Artificial intelligence for mental health: a narrative review of applications, challenges, and future directions in digital health. Digit Health. 2025;11:20552076251395548. 10.1177/20552076251395548 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Cruz‐Gonzalez P, He AWJ, Lam EP, Ng IMC, Li MW, Hou R, et al. Artificial intelligence in mental health care: a systematic review of diagnosis, monitoring, and intervention applications. Psychol Med. 2025;55:e18. 10.1017/S0033291724003295 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Vasilchenko KF, Chumakov EM. Current status, challenges and future prospects in computational psychiatry: a narrative review. Consort Psychiatr. 2023;4(3):33–42. 10.17816/CP11244 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Stein DJ, Shoptaw SJ, Vigo DV, Lund C, Cuijpers P, Bantjes J, et al. Psychiatric diagnosis and treatment in the 21st century: paradigm shifts versus incremental integration. World Psychiatry. 2022;21(3):393–414. 10.1002/wps.20998 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Chen ZS, Kulkarni P (Param), Galatzer‐Levy IR, Bigio B, Nasca C, Zhang Y. Modern views of machine learning for precision psychiatry. Patterns. 2022;3(11):100602. 10.1016/j.patter.2022.100602 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Chen ZS, Schultebraucks K, Wu W. A cautionary tale for AI and machine learning in psychiatry. Transl Psychiatry. 2026;16(1):136. 10.1038/s41398-026-03930-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Huys QJM, Maia TV, Frank MJ. Computational psychiatry as a bridge from neuroscience to clinical applications. Nat Neurosci. 2016;19(3):404–413. 10.1038/nn.4238 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Fernandes BS, Williams LM, Steiner J, Leboyer M, Carvalho AF, Berk M. The new field of ‘precision psychiatry’. BMC Med. 2017;15(1):80. 10.1186/s12916-017-0849-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Torous J, Linardon J, Goldberg SB, Sun S, Bell I, Nicholas J, et al. The evolving field of digital mental health: current evidence and implementation issues for smartphone apps, generative artificial intelligence, and virtual reality. World Psychiatry. 2025;24(2):156–174. 10.1002/wps.21299 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Hua Y, Na H, Li Z, Liu F, Fang X, Clifton D, et al. A scoping review of large language models for generative tasks in mental health care. npj Digit Med. 2025;8(1):230. 10.1038/s41746-025-01611-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Jin Y, Liu J, Li P, Wang B, Yan Y, Zhang H, et al. The applications of large language models in mental health: scoping review. J Med Internet Res. 2025;27:e69284. 10.2196/69284 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Colby KM, Weber S, Hilf FD. Artificial paranoia. Artif Intell. 1971;2(1):1–25. 10.1016/0004-3702(71)90002-6 [DOI] [Google Scholar]
  • 13. Colby KM, Hilf FD, Weber S, Kraemer HC. Turing‐like indistinguishability tests for the validation of a computer simulation of paranoid processes. Artif Intell. 1972;3:199–221. 10.1016/0004-3702(72)90049-5 [DOI] [Google Scholar]
  • 14. Overby MA. Psyxpert: an expert system prototype for aiding psychiatrists in the diagnosis of psychotic disorders. Comput Biol Med. 1987;17(6):383–393. 10.1016/0010-4825(87)90056-4 [DOI] [PubMed] [Google Scholar]
  • 15. Clarke DM, Smith GC, Herrman HE, Mckenzie DP. Monash Interview for Liaison Psychiatry (MILP) development, reliability, and procedural validity. Psychosomatics. 1998;39(4):318–328. 10.1016/S0033-3182(98)71320-9 [DOI] [PubMed] [Google Scholar]
  • 16. Cohen JD, Servan‐Schreiber D. Context, cortex, and dopamine: a connectionist approach to behavior and biology in schizophrenia. Psychol Rev. 1992;99(1):45–77. 10.1037/0033-295X.99.1.45 [DOI] [PubMed] [Google Scholar]
  • 17. Braver TS, Barch DM, Cohen JD. Cognition and control in schizophrenia: a computational model of dopamine and prefrontal function. Biol Psychiatry. 1999;46(3):312–328. 10.1016/S0006-3223(99)00116-X [DOI] [PubMed] [Google Scholar]
  • 18. Hoffman RE. Neural network simulations, cortical connectivity, and schizophrenic psychosis. MD Comput Comput Med Pract. 1997;14(3):200–208. PubMed PMID: 9151510. [PubMed] [Google Scholar]
  • 19. Hoffman RE, McGlashan TH. Parallel distributed processing and the emergence of schizophrenic symptoms. Schizophr Bull. 1993;19(1):119–140. 10.1093/schbul/19.1.119 [DOI] [PubMed] [Google Scholar]
  • 20. Davatzikos C, Shen D, Gur RC, Wu X, Liu D, Fan Y, et al. Whole‐brain morphometric study of schizophrenia revealing a spatially complex set of focal abnormalities. Arch Gen Psychiatry. 2005;62(11):1218. 10.1001/archpsyc.62.11.1218 [DOI] [PubMed] [Google Scholar]
  • 21. Kambeitz J, Cabral C, Sacchet MD, Gotlib IH, Zahn R, Serpa MH, et al. Detecting neuroimaging biomarkers for depression: a meta‐analysis of multivariate pattern recognition studies. Biol Psychiatry. 2017;82(5):330–338. 10.1016/j.biopsych.2016.10.028 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Orrù G, Pettersson‐Yeo W, Marquand AF, Sartori G, Mechelli A. Using Support Vector Machine to identify imaging biomarkers of neurological and psychiatric disease: a critical review. Neurosci Biobehav Rev. 2012;36(4):1140–1152. 10.1016/j.neubiorev.2012.01.004 [DOI] [PubMed] [Google Scholar]
  • 23. Nielsen AN, Barch DM, Petersen SE, Schlaggar BL, Greene DJ. Machine learning with neuroimaging: evaluating its applications in psychiatry. Biol Psychiatry Cogn Neurosci Neuroimaging. 2020;5(8):791–798. 10.1016/j.bpsc.2019.11.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Montague PR, Dolan RJ, Friston KJ, Dayan P. Computational psychiatry. Trends Cognit Sci. 2012;16(1):72–80. 10.1016/j.tics.2011.11.018 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Wang XJ, Krystal JH. Computational psychiatry. Neuron. 2014;84(3):638–654. 10.1016/j.neuron.2014.10.018 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Friston KJ, Stephan KE, Montague R, Dolan RJ. Computational psychiatry: the brain as a phantastic organ. Lancet Psychiatry. 2014;1(2):148–158. 10.1016/S2215-0366(14)70275-5 [DOI] [PubMed] [Google Scholar]
  • 27. Tran T, Luo W, Phung D, Harvey R, Berk M, Kennedy RL, et al. Risk stratification using data from electronic medical records better predicts suicide risks than clinician assessments. BMC Psychiatry. 2014;14(1):76. 10.1186/1471-244X-14-76 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. van Dinteren R, Arns M, Kenemans L, Jongsma MLA, Kessels RPC, Fitzgerald P, et al. Utility of event‐related potentials in predicting antidepressant treatment response: an iSPOT‐D report. Eur Neuropsychopharmacol. 2015;25(11):1981–1990. 10.1016/j.euroneuro.2015.07.022 [DOI] [PubMed] [Google Scholar]
  • 29. Widge AS, Bilge MT, Montana R, Chang W, Rodriguez CI, Deckersbach T, et al. Electroencephalographic biomarkers for treatment response prediction in major depressive illness: a meta‐analysis. Am J Psychiatry. 2019;176(1):44–56. 10.1176/appi.ajp.2018.17121358 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Chekroud AM, Zotti RJ, Shehzad Z, Gueorguieva R, Johnson MK, Trivedi MH, et al. Cross‐trial prediction of treatment outcome in depression: a machine learning approach. Lancet Psychiatry. 2016;3(3):243–250. 10.1016/S2215-0366(15)00471-X [DOI] [PubMed] [Google Scholar]
  • 31. Barak‐Corren Y, Castro VM, Nock MK, Mandl KD, Madsen EM, Seiger A, et al. Validation of an electronic health record–based suicide risk prediction modeling approach across multiple health care systems. JAMA Netw Open. 2020;3(3):e201262. 10.1001/jamanetworkopen.2020.1262 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Salazar De Pablo G, Studerus E, Vaquerizo‐Serrano J, Irving J, Catalan A, Oliver D, et al. Implementing precision psychiatry: a systematic review of individualized prediction models for clinical practice. Schizophr Bull. 2021;47(2):284–297. 10.1093/schbul/sbaa120 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Onnela JP, Rauch SL. Harnessing smartphone‐based digital phenotyping to enhance behavioral and mental health. Neuropsychopharmacology. 2016;41(7):1691–1696. 10.1038/npp.2016.7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Huckvale K, Venkatesh S, Christensen H. Toward clinical digital phenotyping: a timely opportunity to consider purpose, quality, and safety. npj Digit Med. 2019;2(1):88. 10.1038/s41746-019-0166-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Bufano P, Laurino M, Said S, Tognetti A, Menicucci D. Digital phenotyping for monitoring mental disorders: systematic review. J Med Internet Res. 2023;25:e46778. 10.2196/46778 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Khoo LS, Lim MK, Chong CY, McNaney R. Machine learning for multimodal mental health detection: a systematic review of passive sensing approaches. Sensors. 2024;24(2):348. 10.3390/s24020348 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Garriga R, Mas J, Abraha S, Nolan J, Harrison O, Tadros G, et al. Machine learning model to predict mental health crises from electronic health records. Nat Med. 2022;28(6):1240–1248. 10.1038/s41591-022-01811-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Guerreiro J, Garriga R, Lozano Bagén T, Sharma B, Karnik NS, Matić A. Transatlantic transferability and replicability of machine‐learning algorithms to predict mental health crises. npj Digit Med. 2024;7(1):227. 10.1038/s41746-024-01203-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Center for Devices and Radiological Health . Artificial intelligence‐enabled device software functions: lifecycle management and marketing submission recommendations [Internet]. FDA; 2025. [cited 2026 Mar 17]. Available from: https://www.fda.gov/regulatory‐information/search‐fda‐guidance‐documents/artificial‐intelligence‐enabled‐device‐software‐functions‐lifecycle‐management‐and‐marketing [Google Scholar]
  • 40. Wang YF, Li MD, Wang SH, Fang Y, Sun J, Lu L, et al. Large language models in clinical psychiatry: applications and optimization strategies. World J Psychiatr. 2025;15(11). 10.5498/wjp.v15.i11.108199 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Artsi Y, Sorin V, Glicksberg BS, Korfiatis P, Nadkarni GN, Klang E. Large language models in real‐world clinical workflows: a systematic review of applications and implementation. Front Digit Health. 2025;7:1659134. 10.3389/fdgth.2025.1659134 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Shumate JN, Rozenblit E, Flathers M, Larrauri CA, Hau C, Xia W, et al. Governing AI in mental health: 50‐state legislative review. JMIR Ment Health. 2025;12:e80739. 10.2196/80739 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Martinez‐Martin N, Greely HT, Cho MK. Ethical development of digital phenotyping tools for mental health applications: Delphi study. JMIR mHealth uHealth. 2021;9(7):e27343. 10.2196/27343 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Fischer BA. A review of American psychiatry through its diagnoses: the history and development of the diagnostic and statistical manual of mental disorders. J Nerv Ment Dis. 2012;200(12):1022–1030. 10.1097/NMD.0b013e318275cf19 [DOI] [PubMed] [Google Scholar]
  • 45. Shorter E. The history of nosology and the rise of the Diagnostic and Statistical Manual of Mental Disorders . Dialogues Clin Neurosci. 2015;17(1):59–67. 10.31887/DCNS.2015.17.1/eshorter [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Kulikowski CA. Beginnings of artificial intelligence in medicine (AIM): computational artifice assisting scientific inquiry and clinical art – with reflections on present AIM challenges. Yearb Med Inform. 2019;28(1):249–256. 10.1055/s-0039-1677895 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Feigenbaum EA. Expert systems: looking back and looking ahead. In: Wilhelm R, editor. GI – 10. Jahrestagung [Internet] (Brauer W, editor. Informatik‐Fachberichte). Berlin: Springer Berlin Heidelberg; 1980. [cited 2026 Jun 3]. p. 1–14. 10.1007/978-3-642-67838-7_1 [DOI] [Google Scholar]
  • 48. Valton V, Romaniuk L, Douglas Steele J, Lawrie S, Seriès P. Comprehensive review: computational modelling of schizophrenia. Neurosci Biobehav Rev. 2017;83:631–646. 10.1016/j.neubiorev.2017.08.022 [DOI] [PubMed] [Google Scholar]
  • 49. Kambeitz J, Kambeitz‐Ilankovic L, Leucht S, Wood S, Davatzikos C, Malchow B, et al. Detecting neuroimaging biomarkers for schizophrenia: a meta‐analysis of multivariate pattern recognition studies. Neuropsychopharmacology. 2015;40(7):1742–1751. 10.1038/npp.2015.22 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Walker RL, Shortreed SM, Ziebell RA, Johnson E, Boggs JM, Lynch FL, et al. Evaluation of electronic health record‐based suicide risk prediction models on contemporary data. Appl Clin Inf. 2021;12(4):778–787. 10.1055/s-0041-1733908 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Chen L, Xia C, Sun H. Recent advances of deep learning in psychiatric disorders. Precis Clin Med. 2020;3(3):202–213. 10.1093/pcmedi/pbaa029 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Marquand AF, Kia SM, Zabihi M, Wolfers T, Buitelaar JK, Beckmann CF. Conceptualizing mental disorders as deviations from normative functioning. Mol Psychiatr. 2019;24(10):1415–1424. 10.1038/s41380-019-0441-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Xu Y, Fang Z, Lin W, Jiang Y, Jin W, Balaji P, et al. Evaluation of large language models on mental health: from knowledge test to illness diagnosis. Front Psychiatr. 2025;16:1646974. 10.3389/fpsyt.2025.1646974 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Jackson RG, Patel R, Jayatilleke N, Kolliakou A, Ball M, Gorrell G, et al. Natural language processing to extract symptoms of severe mental illness from clinical text: the Clinical Record Interactive Search Comprehensive Data Extraction (CRIS‐CODE) project. BMJ Open. 2017;7(1):e012012. 10.1136/bmjopen-2016-012012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Newby D, Taylor N, Joyce DW, Winchester LM. Optimising the use of electronic medical records for large scale research in psychiatry. Transl Psychiatry. 2024;14(1):232. 10.1038/s41398-024-02911-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Dikaios K, Rempel S, Dumpala SH, Oore S, Kiefte M, Uher R. Applications of speech analysis in psychiatry. Harv Rev Psychiatr. 2023;31(1):1–13. 10.1097/HRP.0000000000000356 [DOI] [PubMed] [Google Scholar]
  • 57. Kappen M, Vanderhasselt MA, Slavich GM. Speech as a promising biosignal in precision psychiatry. Neurosci Biobehav Rev. 2023;148:105121. 10.1016/j.neubiorev.2023.105121 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Chancellor S, De Choudhury M. Methods in predictive techniques for mental health status on social media: a critical review. npj Digit Med. 2020;3(1):43. 10.1038/s41746-020-0233-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Guntuku SC, Yaden DB, Kern ML, Ungar LH, Eichstaedt JC. Detecting depression and mental illness on social media: an integrative review. Curr Opin Behav Sci. 2017;18:43–49. 10.1016/j.cobeha.2017.07.005 [DOI] [Google Scholar]
  • 60. Torous J, Bucci S, Bell IH, Kessing LV, Faurholt‐Jepsen M, Whelan P, et al. The growing field of digital psychiatry: current evidence and the future of apps, social media, chatbots, and virtual reality. World Psychiatry. 2021;20(3):318–335. 10.1002/wps.20883 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Philippe TJ, Sikder N, Jackson A, Koblanski ME, Liow E, Pilarinos A, et al. Digital health interventions for delivery of mental health care: systematic and comprehensive meta‐review. JMIR Ment Health. 2022;9(5):e35159. 10.2196/35159 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Liu X, Rivera SC, Moher D, Calvert MJ, Denniston AK. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT‐AI extension. BMJ. 2020:m3164. 10.1136/bmj.m3164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63. Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, et al. Reporting guideline for the early‐stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE‐AI. Nat Med. 2022;28(5):924–933. 10.1038/s41591-022-01772-9 [DOI] [PubMed] [Google Scholar]
  • 64. Reading Turchioe M, Desai P, Harkins S, Kim J, Kumar S, Zhang Y, et al. Differing perspectives on artificial intelligence in mental healthcare among patients: a cross‐sectional survey study. Front Digit Health. 2024;6:1410758. 10.3389/fdgth.2024.1410758 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Celi LA, Cellini J, Charpignon ML, Dee EC, Dernoncourt F, Eber R, et al. Sources of bias in artificial intelligence that perpetuate healthcare disparities—a global review. PLOS Digit Health. 2022;1(3):e0000022. 10.1371/journal.pdig.0000022 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Muehlematter UJ, Daniore P, Vokinger KN. Approval of artificial intelligence and machine learning‐based medical devices in the USA and Europe (2015–20): a comparative analysis. Lancet Digit Health. 2021;3(3):e195–e203. 10.1016/S2589-7500(20)30292-2 [DOI] [PubMed] [Google Scholar]
  • 67. Meskó B, Topol EJ. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digit Med. 2023;6(1):120. 10.1038/s41746-023-00873-0 [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Psychiatric Research and Clinical Practice are provided here courtesy of Wiley

RESOURCES