Abstract
Currently, artificial intelligence (AI) is clinically relevant to mood and anxiety care, but the evidence base is uneven across use cases. This narrative review synthesizes recent literature most relevant to clinicians and investigators. Five themes dominate the current field: patient-facing adjunctive tools, failure modes and safety risks, clinician-facing decision support, passive sensing and measurement infrastructure, and governance. Recent randomized evidence supports a narrow efficacy claim for structured chatbot interventions, with small improvements in depressive and anxiety symptoms and more consistent effects on engagement than on symptom superiority. These studies do not support autonomous psychotherapy, and they do not establish a therapeutic advantage for open-ended large language model systems over more constrained designs. Safety studies, by contrast, identify active concerns: harmful endorsement, weak youth risk assessment, inconsistent crisis handling, and anxiety/OCD reassurance loops. The strongest current clinical signal lies in supervised clinician-facing decision support, where recent trials of AI-assisted antidepressant selection improved treatment persistence and some downstream symptom outcomes. Passive sensing detects behaviorally meaningful signals, but evidence that alert-driven deployment improves care remains insufficient for routine practice. Across stakeholder and policy sources, the most defensible deployment model is human-in-the-loop, stepped, and bounded by explicit handoff rules.
Keywords: Artificial intelligence, Anxiety, Depression, Chatbot, Phenotyping, Governance, Implementation
1. Introduction
Artificial intelligence is already part of the lived ecology of mood and anxiety care. In a nationally representative US survey, 13.1% of adolescents and young adults reported using generative AI for mental health advice; among users, 65.5% sought such advice at least monthly, 92.7% rated it as somewhat or very helpful, and use was highest among those aged 18–21 years (22.2%) [1]. For clinicians, this adoption curve changes the relevant question. The key issue is no longer whether AI will enter mood and anxiety care. It already has. The decisive question is where AI adds value without degrading safety, alliance, privacy, or judgment [1], [2], [3], [4], [5], [6], [7].
This narrative review focuses on applications most relevant to clinicians and investigators. We organize the literature into five themes: patient-facing adjunctive tools, failure modes in higher-risk settings, clinician-facing decision support, passive measurement infrastructure, and governance. Throughout, we separate association, prediction, and intervention. That distinction is not semantic. A chatbot can increase engagement without improving symptoms, a predictive model can explain variance without changing outcomes, and a mechanistic hypothesis can be conceptually attractive without being clinically validated [8], [9], [10], [11], [12], [13], [14], [15], [16], [17], [18], [19], [20]. This article is a focused narrative review of recent systematic reviews, randomized trials, simulation-based safety studies, qualitative stakeholder work, and professional guidance documents with direct relevance to mood and anxiety care. To emphasize timeliness and decision utility, the review prioritizes evidence published during 2024–2026 and asks three practical questions: what can be used now, what warrants guarded pilots, and what should remain outside unsupervised AI deployment.
2. An inferential framework for current AI in mental health
Three analytic errors recur in this literature. First, engagement is often treated as a proxy for efficacy. Second, transcript quality or apparent empathy is treated as a proxy for safety. Third, prediction is treated as if it were causal or mechanistic explanation. These errors matter because they invite premature clinical deployment [10], [11], [12], [13], [14], [15], [16], [17], [18], [19], [20]. For practical purposes, current AI applications in mood and anxiety care fall into three lanes. The first is patient-facing adjuncts, including chatbots, coaching tools, and between-session supports. The second is clinician-facing support, including treatment selection and structured workflow augmentation. The third is measurement infrastructure, including passive sensing and state prediction. The evidence standard should differ across these lanes. Patient-facing tools require randomized evidence and safety data. Clinician-facing tools require incremental value over clinician judgment within supervised care. Passive sensing requires prospective proof that acting on alerts improves outcomes rather than merely generating additional signals [13], [14], [15], [16], [17], [18], [19].
Table 1 summarizes the evidence class, main clinical signal, inferential limit, and current ADAA-relevant action for each theme.
Table 1.
Current themes in AI for mental health care and their present clinical status.
| Theme | Current evidence | Main inferential limit | ADAA-relevant action |
|---|---|---|---|
| Patient-facing adjuncts | Small symptomatic benefits for structured chatbot interventions; improved engagement for some generative-AI-enhanced CBT tools [8], [10], [11] | High heterogeneity, high risk of bias, limited safety reporting, and no convincing superiority of open-ended LLM chatbots [8], [9] | Use as adjunctive between-session support, psychoeducation, or adherence coaching |
| Failure modes | Unsafe endorsement is possible; youth risk assessment is weak; crisis responses remain variable; reassurance loops may worsen anxiety/OCD [21], [22], [23], [24] | Most studies do not estimate real-world incidence, but they clearly establish failure potential [21], [22], [23], [24] | Do not delegate suicidality, youth safeguarding, ERP/OCD reassurance, or mania/psychosis management |
| Clinician-facing decision support | Supervised treatment selection can improve persistence and some downstream outcomes; modest prediction can inform shared decisions [13], [14], [15] | Open-label designs, missingness, interpretability limits, and modest variance explained [13], [14], [15] | Prioritize reviewable decision support and workflow augmentation |
| Passive sensing | Mobility, sleep, activity, phone use, and physiology carry symptom-related signals [16], [17], [18] | Mostly associative evidence; thin intervention evidence; false positives, privacy concerns, and weak external validation [16], [17], [18] | Restrict to monitored pilots; require person-specific baselines and prospective utility |
| Governance and equity | Patients prefer humans in the loop; lifecycle governance and postmarket monitoring are necessary [2], [3], [4], [5], [6], [7] | Guidance is implementation-oriented rather than proof of efficacy [3], [4], [5], [6], [7] | Require disclosure, subgroup reporting, audit trails, drift monitoring, and handoff rules |
Abbreviations: CBT, cognitive-behavioral therapy; ERP, exposure and response prevention; LLM, large language model.
3. Patient-facing adjuncts: a real efficacy signal, but a narrow one
A 2026 meta-analysis of 39 randomized trials provides the strongest broad evidence for chatbot interventions in depressive and anxiety symptoms [8]. Pooled post-intervention effects were small but statistically significant for depression (Hedges g = 0.31, 95% CI 0.17–0.46) and anxiety (g = 0.28, 95% CI 0.05–0.51). The effect on depressive symptoms was larger in clinical and subclinical samples than in nonclinical samples. That is an efficacy signal but it is not a validation of autonomous psychotherapy. Heterogeneity was high for both outcomes, most studies were at high overall risk of bias, most outcomes were self-reported, and safety monitoring was underdeveloped [8]. The clinical conclusion is therefore limited: structured chatbot interventions can reduce symptoms modestly in some settings, but the magnitude and reliability of benefit remain constrained. The model-class question is also frequently misframed. A 2025 systematic review and meta-analysis comparing rule-based with LLM-based chatbots concluded that the evidence remains stronger for more constrained systems and that robust evidence of superiority for LLM-based chatbots is lacking [9]. Clinically, generative capability is not itself a therapeutic mechanism. In mental health, behavioral structure, scope constraints, and intervention design still appear to matter more than conversational fluency alone [8], [9].
Two recent trials further emphasize this point. In a 6-week randomized trial, a generative-AI-enhanced CBT app substantially increased user engagement, including more frequent use and longer interaction time, but symptom and safety outcomes were similar to a digital workbook comparator [10]. The plausible near-term value of generative systems may therefore lie in adherence support rather than robust and clinically reliable autonomous symptom reduction. In a separate randomized trial, richer social cues improved PHQ-9 and GAD-7 outcomes, adherence, satisfaction, and therapeutic alliance relative to text-only chatbot delivery [11]. Taken together, these studies suggest that some of what is being labeled an AI effect is better understood as a design effect. A warmer interpersonal wrapper, better conversational pacing, and more usable homework support may improve engagement and possibly outcomes, but those gains should not be confused with evidence that open-ended consumer chatbots can safely deliver therapy [10], [11].
In a 2026 study, a psychotherapy-specific cognitive layer improved expert-rated CBT competence relative to standalone LLMs (mean Cognitive Therapy Rating Scale score 4.53 vs 3.16), and higher activation of that layer in observational analyses was associated with better transcript quality and better symptom trajectories [12]. These findings are promising, but they do not generalize to generic consumer chatbots. A domain-specific architecture with explicit safety and clinical reasoning layers is a different intervention class from a general-purpose LLM prompt wrapper [12]. For a practicing clinician, the defensible near-term role of patient-facing AI is therefore narrow: low-intensity adjunctive support, between-session CBT reinforcement, psychoeducation, and adherence coaching for appropriately selected patients. The current evidence does not justify treating unsupervised AI as a substitute for clinician-delivered CBT, exposure and response prevention, crisis assessment, or complex case management [8], [9], [10], [11], [12].
4. Failure modes in mood and anxiety care: why risk is not an edge case
The safety literature is less mature than the efficacy literature but clinically important. A simulation-based study of 10 publicly available therapy and companion chatbots found explicit endorsement of harmful or ill-advised adolescent proposals in 19 of 60 opportunities; no bot reliably opposed all dangerous suggestions [21]. Companion bots performed worst. This design does not estimate real-world incidence, but it does establish something more basic: unsafe limit-setting remains technically possible in routine consumer systems [21]. A separate 2025 cross-sectional evaluation of youth-facing generative-AI psychotherapy chatbots found high ratings for accessibility and frequent avoidance of overt misinformation, but only 31% of therapeutic-approach ratings and 39% of monitoring or assessment-of-risk ratings were judged high quality [22]. This pattern is clinically familiar: surface fluency can coexist with poor method. For youth care, acceptable usability is not a sufficient safety marker [22].
Crisis-response behavior has improved, but inconsistency remains. In a two-phase content analysis of suicide-related prompts, later-generation chatbots were more likely to mention 988, address lethality, and encourage urgent help-seeking than earlier systems [23]. That improvement is welcome. It is not sufficient grounds to delegate suicide triage to AI. Response quality varied across models, prompts were framed around concern for another person rather than first-person suicidal intent, and one chatbot generated an unsolicited suicide-related image [23].
For anxiety disorders and OCD, a further problem is that apparently supportive responses can still be counterproductive. A recent conceptual model argues that general-purpose chatbots may reinforce reassurance seeking, intolerance of uncertainty, perfectionism, re-checking, and avoidance - the very processes that maintain OCD and many anxiety disorders [24]. The APA health advisory echoes this concern, warning that generative-AI chatbots and wellness apps may not have adequate validation or safety design for mental health use and calling for explicit guardrails [4]. For clinicians, this is not a peripheral issue. In exposure-based treatment, a system that reliably reduces uncertainty on demand can work directly against treatment goals [4], [24]. Certain domains should therefore remain outside unsupervised AI delegation: suicide and self-harm assessment, youth safeguarding, abuse or exploitation disclosures, OCD reassurance and exposure-related decision-making, mania or psychosis, and any scenario that requires firm clinical limit-setting [4], [21], [22], [23], [24]. Minimum safeguards for any patient-facing deployment should include explicit disclosure of AI use, validated detection of self-harm language, direct routing to crisis resources, disorder-specific cautions for anxiety and OCD, age-appropriate safeguards, and auditable human handoff thresholds [3], [4], [5], [6], [7].
5. Clinician-facing AI: where the current clinical signal is strongest
Among current applications, the most credible near-term role for AI in mood and anxiety care is not autonomous psychotherapy but supervised clinical decision support. In the PETRUSHKA randomized clinical trial, a web-based tool that combined clinical and demographic predictors with patient preferences reduced antidepressant discontinuation at 8 weeks from 27% under usual care to 17% with decision support (adjusted relative risk 0.62, 95% CI 0.44–0.88) [13]. By 24 weeks, PETRUSHKA was also associated with lower PHQ-9 and GAD-7 scores (adjusted mean differences −1.92 and −1.39, respectively) [13]. These findings matter because the primary endpoint was treatment persistence. The limitation is equally important: the study was open-label for patients and clinicians, secondary outcomes had substantial missingness, and the model remained only partially interpretable [13].
The AID-ME cluster randomized trial likewise suggests that supervised AI-based treatment selection may improve outcomes in major depression, with remission observed in 28.6% of analyzed patients in the AI-CDSS group versus none in active control, and no serious adverse events attributed to the system [14]. The inferential limit is straightforward: these are encouraging results for supervised care, not a basis for autonomous prescribing. The clinical question is whether decision support adds value beyond severity assessment, comorbidity, patient preference, and clinician judgment. That is the correct benchmark for this class of tools [13], [14].
Predictive treatment selection remains modest rather than deterministic. In a prognostic study of 883 patients, an elastic net model using 27 mostly self-reported predictors explained about 19% of the variance in improvement for internet-delivered CBT in unseen data and showed clearer treatment specificity when retrained on a single-treatment cohort [15]. That is useful signal, but it also means that most clinically relevant variation remained unexplained [15]. Therefore, the utility of AI clinical prediction tools should not be overinterpreted. A model that explains part of the variance can support shared decisions; it cannot replace formulation.
In a clinical setting, clinician-facing AI should therefore be prioritized where outputs are structured, reviewable, and clinically bounded. Candidate uses include ranked treatment options, pre-visit synthesis, standardized intake support, symptom questionnaire integration, and referral navigation. Although these workflow applications remain less rigorously tested than treatment-selection tools, they are consistent with stakeholder preferences: patients with anxiety were more comfortable with administrative, screening, and lower-risk supplemental uses than with AI-only care [2]. The practical aim should be to reduce cognitive and administrative load so clinicians can spend more time on alliance, context, and judgment rather than to accelerate automation for its own sake [2], [3], [13].
Before adopting clinician-facing AI, mood and anxiety programs should ask four questions. Does the tool improve a patient-centered endpoint rather than only an internal model metric? Does it demonstrate incremental value over current practice? Is performance calibrated and subgroup-stratified? And can a clinician readily review, override, and document the recommendation? If the answer to any of these is no, the tool is not deployment-ready [3], [6], [7], [13], [14], [15].
6. Passive sensing and digital phenotyping: promising infrastructure, insufficient intervention evidence
Passive sensing literature shows that smartphones and wearables can capture behaviorally meaningful signals related to mood and anxiety symptoms. A 2024 systematic review of digital phenotyping in nonclinical adults found repeated associations between symptoms and mobility, sleep regularity, physical activity, phone use, and social interaction; 78% of the included studies used machine-learning methods [16]. A 2025 scoping review of passive sensing across mental disorders likewise identified recurring signal domains, including sleep, activity, physiology, and social behavior [18].
However, there is still a gap from signal detection to clinical utility. When the literature is restricted to longitudinal depression monitoring and prediction of clinical states, only 9 studies met criteria in one 2025 systematic review, 6 reported some prediction capability, and only 1 distinguished worsening, relapse, or recovery states [17]. Across reviews, common limitations include small samples, short monitoring windows, high attrition, self-reported labels, inconsistent privacy reporting, and weak external validation [16], [17], [18]. The same sensor feature can also mean different things across people and contexts. Reduced mobility may reflect depression, but it may also reflect remote work, caregiving, illness, or adaptive rest [16].
For clinicians, this means passive data should currently be treated as research or guarded-pilot infrastructure rather than routine care. A clinically actionable alert should meet a higher threshold than statistical detectability. At minimum, it should be anchored to a person-specific baseline, handle missingness explicitly, provide an interpretable rationale, have a manageable false-positive burden, and show prospective evidence that acting on the alert improves outcomes or care processes [16], [17], [18]. Without that evidence, passive sensing risks becoming noisy surveillance rather than measurement-based care.
7. Governance, equity, and human oversight: the deployment question is now central
Stakeholder data and professional guidance converge on a common operational principle: AI in mental health should be human-in-the-loop by default. In qualitative interviews, adults with mild to moderate anxiety saw potential benefits in access and convenience but expressed concerns about empathy, privacy, technical limitations, and replacement of therapy; most preferred some degree of human involvement [2]. These views are consistent with APA ethical guidance for health service psychology and the APA advisory on generative-AI chatbots and wellness apps, both of which emphasize guardrails rather than unrestricted substitution [3], [4].
Recent policy documents reinforce the same direction. WHO guidance on large multi-modal models in health emphasizes ethics and governance rather than unqualified adoption [5]. FDA's 2025 draft guidance on AI-enabled device software functions places lifecycle risk management at the center of premarket and postmarket evaluation [6]. The FDA Digital Health Advisory Committee's 2025 discussion of generative-AI-enabled digital mental health medical devices likewise focused on premarket evidence, limitations of use, and postmarket monitoring [7]. This is the correct regulatory logic for mental health, where task complexity, conversational ambiguity, and risk heterogeneity are unusually high.
For mental health systems, a clinically credible governance architecture should include at least six elements. First, disclose when AI is used, what data it accesses, and how patients can opt out. Second, report performance across relevant subgroups, including age, diagnosis, symptom severity, language, and care setting. Third, maintain audit trails for prompts, outputs, clinician edits, and handoffs. Fourth, monitor model drift and re-evaluate high-risk workflows after product updates. Fifth, red-team disorder-specific failure modes, particularly suicidality, OCD reassurance, youth safety, mania, exploitation, and hallucinated clinical advice. Sixth, define explicit thresholds for human takeover rather than leaving escalation to discretionary impression [2], [3], [4], [5], [6], [7].
Equity concerns deserve special emphasis. If helpfulness, engagement, or language style differ across demographic groups, a mental health AI product may widen disparities even when average performance appears acceptable. The McBain survey already suggests differential perceived helpfulness across racial groups among youth users [1]. In high-stakes mental health deployment, average performance is an insufficient fairness metric [1], [2], [3], [4], [5], [6], [7].
A practical stepped-care boundary map for AI tasks in mood and anxiety care is summarized in Table 2.
Table 2.
Stepped-care boundaries for AI tasks in mood and anxiety disorders.
| Deployment tier | Example tasks | Current status | Conditions for use |
|---|---|---|---|
| Low-risk automation | Scheduling, psychoeducation, reminders, questionnaire administration, basic navigation | Suitable now | Disclose AI use; no high-risk clinical advice; easy opt-out |
| Adjunctive between-session support | CBT homework reinforcement, symptom check-ins, adherence prompts | Suitable in selected patients | Structured content, monitored worsening, explicit human fallback |
| Supervised clinician support | Treatment ranking, pre-visit synthesis, standardized intake, referral matching | Reasonable current priority | Clinician review, calibration, auditability, subgroup checks |
| Passive sensing and alerting | Relapse-risk alerts, behavior-change flags, monitoring dashboards | Research or guarded pilot only | Person-specific baselines, false-positive control, prospective utility |
| Autonomous high-risk care | Crisis response, youth safeguarding, OCD reassurance decisions, mania or psychosis management | Not supported | Retain human clinical responsibility |
Abbreviations: AI, artificial intelligence; CBT, cognitive-behavioral therapy; OCD, obsessive-compulsive disorder. This stepped-care map is synthesized from the clinical and governance literature reviewed in [2], [3], [4], [5], [6], [7], [8], [9], [10], [11], [12], [13], [14], [15], [16], [17], [18], [19], [20], [21], [22].
8. A translational research agenda: from prediction to mechanism and from novelty to clinical utility
Current AI in mental health is dominated by operational and predictive applications. That is not necessarily a weakness, but it should be made explicit. The precision psychiatry literature argues that clinically useful models must move beyond simple association toward causal prediction under hypothetical interventions [19]. Mechanistic frameworks such as predictive coding offer one candidate language for linking neural-system determinants, symptom expression, and treatment response [20]. However, the studies reviewed here do not yet support strong mechanistic claims for routine anxiety and depression care. Most current tools estimate symptom change, engagement, or selection benefit; they do not identify validated causal mechanisms at the individual level [13], [14], [15], [16], [17], [18], [19], [20].
From these findings one can proposed four research priorities. First, disorder-specific rather than generic design should become the norm, particularly for anxiety disorders, OCD, and youth care [22], [23], [24]. Second, AI trials should separate engagement endpoints from clinical endpoints and include systematic adverse-event monitoring [8], [10], [11]. Third, clinician-facing models should emphasize calibration, external validation, subgroup performance, and incremental value over usual care rather than only discrimination statistics [13], [14], [15], [19]. Fourth, passive sensing should be evaluated within closed clinical workflows that test whether acting on alerts changes patient outcomes, clinician behavior, or service utilization [16], [17], [18]. To be clinically meaningful, future studies should also specify what decision the model is intended to change: treatment selection, monitoring intensity, between-session support, referral matching, or escalation threshold. Models that do not alter a well-defined clinical decision are unlikely to generate durable utility, regardless of technical sophistication [13], [14], [15], [16], [17], [18], [19].
9. Conclusions
AI in mental health has advanced far enough that undifferentiated statements about the hope or perils of AI are no longer useful. The literature supports a bounded set of claims. Structured chatbots can yield small symptomatic benefits, but current evidence is strongest for constrained adjunctive uses rather than autonomous therapy [8], [9], [10], [11]. The most urgent current risks are unsafe advice, weak crisis handling, and reinforcement of anxiety/OCD maintaining processes [4], [21], [22], [23], [24]. The strongest clinical signal lies in supervised clinician-facing decision support and workflow augmentation, not in replacement of the therapist [13], [14], [15]. Passive sensing remains promising but insufficiently mature for routine alert-driven care [16], [17], [18]. Across stakeholder and policy sources, the most defensible deployment model is human-in-the-loop, stepped, and bounded by explicit handoff rules [2], [3], [4], [5], [6], [7]. For clinicians and investigators, the practical task is not to decide whether AI is good or bad. It is to determine which specific tasks can be improved now, which require guarded pilots, and which should remain firmly within human clinical responsibility (Fig. 1).
Fig. 1.
Author Contributions
Martin P. Paulus is the sole author and is responsible for the conception, drafting, revision, and final approval of the manuscript.
Ethics Approval
Not applicable.
Funding
No specific funding was received for this manuscript.
Conflict of Interest
The author declares no conflict of interest.
Declaration of Competing Interest
The authors report no conflict of interest with respect to the content of this manuscript. Dr. Paulus is an advisor to Spring Care, Inc., a behavioral health startup, he has received royalties for an article about methamphetamine in UpToDate. Dr. Paulus has a consulting agreement with and receives compensation from F. Hoffmann-La Roche Ltd.
Acknowledgements/Funding
This work has been supported in part by The William K. Warren Foundation and the National Institute of General Medical Sciences Center Grant Award Number (1P20GM121312) and the National Institute on Drug Abuse (U01DA050989).
Data Availability
Not applicable.
References
- 1.McBain R.K., Bozick R., Diliberti M., et al. Use of generative AI for mental health advice among US adolescents and young adults. JAMA Netw Open. 2025;8(11) doi: 10.1001/jamanetworkopen.2025.42281. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Lee H.S., Wright C., Ferranto J., et al. Artificial intelligence conversational agents in mental health: patients see potential, but prefer humans in the loop. Front Psychiatry. 2025;15 doi: 10.3389/fpsyt.2024.1505024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.American Psychological Association . American Psychological Association; Washington, DC: 2025. Ethical guidance for AI in the professional practice of health service psychology. [Google Scholar]
- 4.American Psychological Association . American Psychological Association; Washington, DC: 2025. Health advisory on the use of generative AI chatbots and wellness apps for mental health. [Google Scholar]
- 5.World Health Organization . World Health Organization; Geneva: 2025. Ethics and governance of artificial intelligence for health: guidance on large multi-modal models. [Google Scholar]
- 6.U.S. Food and Drug Administration . U.S. Food and Drug Administration; Silver Spring, MD: 2025. Artificial intelligence-enabled device software functions: lifecycle management and marketing submission recommendations. draft guidance for industry and food and drug administration staff. [Google Scholar]
- 7.U.S. Food and Drug Administration . Digital Health Advisory Committee: generative artificial intelligence-enabled digital mental health medical devices. U.S. Food and Drug Administration; Silver Spring, MD: 2025. [Google Scholar]
- 8.Sohn J.S., Ha B.G., Park S., et al. Systematic review and meta-analysis of chatbots in the management of depressive and anxiety symptoms. npj Digit Med. 2026;9:1–17. doi: 10.1038/s41746-026-02566-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Q Du, et al. The efficacy of rule-based versus large language model-based chatbots in alleviating symptoms of depression and anxiety: systematic review and meta-analysis. J Med Internet Res. 2025;27 doi: 10.2196/78186. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.McFadyen J., Habicht J., Dina L.M., et al. Increasing engagement with cognitive-behavioral therapy using generative AI: a randomized controlled trial. Commun Med. 2026;6:13. doi: 10.1038/s43856-025-01321-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Xu S., Ma T. Depression intervention using AI chatbots with social cues: a randomized trial of effectiveness. J Affect Disord. 2025;389 doi: 10.1016/j.jad.2025.119760. [DOI] [PubMed] [Google Scholar]
- 12.Rollwage M., McFadyen J., Juchems K., et al. A cognitive layer architecture to support large-language model performance in psychotherapy interactions. Nat Med. 2026 doi: 10.1038/s41591-026-04278-w. [DOI] [PubMed] [Google Scholar]
- 13.Cipriani A., Fernandes K.B.P., Mulsant B.H., et al. A decision-support system to personalize antidepressant treatment in major depressive disorder: a randomized clinical trial. JAMA. 2026;325:1–13. doi: 10.1001/jama.2026.1327. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Benrimoh D., Whitmore K., Richard M., et al. Artificial intelligence in depression-medication enhancement (AID-ME): a cluster randomized trial of a deep-learning-enabled clinical decision support system for personalized depression treatment selection and management. J Clin Psychiatry. 2025;86(3) doi: 10.4088/JCP.24m15634. [DOI] [PubMed] [Google Scholar]
- 15.Lee C.T., Richards D., Heinzle J., et al. Machine learning model for response to internet-delivered CBT vs antidepressant medication. JAMA Netw Open. 2025;8(11) doi: 10.1001/jamanetworkopen.2025.41639. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Choi A., Ooi A., Lottridge D. Digital phenotyping for stress, anxiety, and mild depression: systematic literature review. JMIR Mhealth Uhealth. 2024;12 doi: 10.2196/40689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Amin R., Schreynemackers S., Oppenheimer H., Petrovic M., Hegerl U., Reich H. Use of mobile sensing data for longitudinal monitoring and prediction of depression severity: systematic review. J Med Internet Res. 2025;27 doi: 10.2196/57418. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Shen S., Qi W., Zeng J., et al. Passive sensing for mental health monitoring using machine learning with wearables and smartphones: scoping review. J Med Internet Res. 2025;27 doi: 10.2196/77066. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Krishnadas R., Leighton S.P., Jones P.B. Precision psychiatry: thinking beyond simple prediction models - enhancing causal predictions. Br J Psychiatry. 2025;226(3):184–188. doi: 10.1192/bjp.2024.258. [DOI] [PubMed] [Google Scholar]
- 20.Shaw A.D., Sumner R.L., Berndt L.C.S. Predictive coding and neurocomputational psychiatry: a mechanistic framework for understanding mental disorders. Front Psychiatry. 2026;16 doi: 10.3389/fpsyt.2025.1713833. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Clark A. The ability of AI therapy bots to set limits with distressed adolescents: simulation-based comparison study. JMIR Ment Health. 2025;12 doi: 10.2196/78414. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Sobowale K., Humphrey D.K., Zhao S.Y. Evaluating generative AI psychotherapy chatbots used by youth: cross-sectional study. JMIR Ment Health. 2025;12 doi: 10.2196/79838. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Campbell L.O., Babb K., Lambie G.W., Hayes B.G. An examination of generative AI response to suicide inquiries: content analysis. JMIR Ment Health. 2025;12 doi: 10.2196/73623. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Golden A., Aboujaoude E. A transdiagnostic model for how general purpose AI chatbots can perpetuate OCD and anxiety disorders. npj Digit Med. 2026;9:1–5. doi: 10.1038/s41746-026-02531-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Not applicable.

