Abstract
Background
The rapid advancement of Large Language Models has sparked heated debate over whether Generative Artificial Intelligence (AI) chatbots can serve as “digital therapists” capable of providing therapeutic support. While much of this discussion focuses on AI’s lack of agency, understood as the absence of mental states, consciousness, autonomy, and intentionality, empirical research on users’ real-world experiences remains limited.
Objective
This study explores how individuals with mental distress experience support from both generative AI chatbots and human psychotherapy in natural and unguided contexts, with a focus on how perceptions of agency shape therapeutic experiences. By drawing on participants’ dual exposure, the study seeks to contribute to the ongoing debate about “AI therapists” by clarifying the role of agency in therapeutic change.
Methods
Sixteen adults who had sought mental health support from both human therapists and ChatGPT participated in semi-structured interviews, during which they shared and compared their experiences with each type of interaction. Transcripts were analyzed using reflexive thematic analysis.
Results
Three themes captured participants’ perceptions of ChatGPT relative to human therapists: (1) encouraging open and authentic self-disclosure but limiting deep exploration; (2) the myth of relationship: caring, acceptance, and understanding; (3) fostering therapeutic change: the promise and pitfalls of data-driven solutions. We propose a conceptual model that illustrates how differences in agency status between AI chatbots and human therapists shape the distinct ways they support individuals with mental distress, with agency functioning as both a strength and a limitation for therapeutic engagement.
Conclusion
Given that agency functions as a double-edged sword in therapeutic interactions, future mental health services should consider integrated care models that combine the non-agential advantages of AI chatbots with the agentic qualities of human therapists. Rather than anthropomorphizing AI chatbots, their non-agential features—such as responsiveness, absence of intentions, objectivity, and disembodiment—should be strategically leveraged to complement specific functions in human-delivered psychotherapy. At the same time, practitioners should maximize the benefits of their agentic qualities while remaining cautious of the risks. The findings should be interpreted with caution as the sample consisted mainly of young, well-educated Chinese participants from a collectivist cultural context, which may limit transferability to other populations, particularly those from individualistic cultures with different mental health literacy levels, stigma patterns, and therapeutic norms.
Clinical trial number
Not applicable.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12888-025-07671-w.
Keywords: Generative Artificial Intelligence, Psychotherapy, Mental distress, Anthropomorphizing, Agency
Background
The mental health crisis is widely recognized as one of the most significant global challenges [1]. Epidemiological data [2] indicated that one in three people would develop a mental health disorder at some point in their lifetime, yet only 36.8% of affected individuals in high-income countries received mental health care, and the rate dropped to 13.7% in lower-middle-income countries [3]. A cross-national survey found that, alongside structural barriers such as the shortage of professionals, attitudinal barriers such as preference for self-reliance, perceived ineffectiveness of treatment, and negative prior experiences also contributed to low treatment rates [4]. To address these challenges, researchers have increasingly explored the use of cutting-edge technologies, such as Artificial Intelligence (AI), to develop alternative mental health solutions.
The evolution of AI mental health chatbots: Can LLMs transform therapeutic support?
AI mental health chatbots are dialogue systems designed to simulate interactive, human-like conversations through text or voice in real time [5], representing a promising technological approach. Recent reviews [6, 7] and meta-analyses studies [8] have shown that AI chatbots hold substantial potential in promoting mental well-being, particularly in reducing symptoms of depression and anxiety. However, most existing evidence concerns rule-based chatbots that rely on predefined conversational flows within frame-based architectures. These systems are typically limited to narrow tasks (e.g., stress reduction, mood tracking) and lack the flexibility needed for complex therapeutic interactions.
The advent of Large Language Models (LLMs) marks a paradigm shift by enabling dynamic, context-aware interactions that approximate human conversation and unlock new possibilities for mental health support. LLMs are large-scale Transformer models trained on extensive datasets [9]. A well-known example is ChatGPT, built on the Transformer decoder architecture [10] and uses self-attention to generate coherent text [11]. It was pretrained on large-scale corpora and later fine-tuned with supervised learning and reinforcement learning from human feedback [12]. These advancements allow ChatGPT to overcome many limitations of rule-based chatbots, such as repetitive, constrained, and unnatural interactions, making it more capable of engaging in open-ended discussions on diverse mental health concerns [8].
Emerging evidence suggests that the performance of AI chatbots may be comparable with well-trained human experts in various psychology-related tasks. For instance, ChatGPT has demonstrated superior performance on social intelligence assessments [13] and generated responses that third-party evaluators rated as more empathetic and helpful than those provided by clinicians [14, 15]. Some researchers further argue that ChatGPT is capable of handling complex clinical tasks, such as generating plausible psychiatric plans [16].
Agency as the central issue in the AI therapist debate
ChatGPT’s promising performance has sparked heated debate about whether LLM-powered AI chatbots (hereafter AI chatbots) could take on more complex therapeutic roles as “digital therapists” [17, 18]. While early findings underscore their transformative potential, are they already capable of offering psychotherapy on their own? Central to this debate is a philosophical and empirical inquiry: can therapeutic change only occur through conversations between two conscious agents [18]?
Therapeutic change refers to a set of beneficial psychological transformations that occur during psychotherapy. It involves not only symptom relief, but also deeper changes in self-understanding, improvements in daily functioning, and enhanced interpersonal relationships [19]. In philosophy of mind, an agent is an entity capable of initiating actions for reasons, where such actions express intentions and are guided by mental states (beliefs, desires, intentions). The core characteristics of an agent are purposeful action (intentionality), self-directed control (autonomy), and reflective awareness (consciousness) [20]. Although recent studies show LLM applications demonstrating greater behavioral flexibility [21], they still lack essential components of agency, such as conscious awareness [22]. Some scholars therefore view AI chatbots as hybrid artifacts with both tool-like and agent-like qualities [18].
This ambiguous status has intensified debate over whether the absence of full agency prevents AI chatbots from serving as digital therapists. Critics argued that they could not enable genuine insight or self-awareness because they were not rational agents capable of interpreting meaning, understanding context, or engaging in reflective dialogue. These processes were considered essential in psychotherapy for individuals to make sense of their inner experiences and integrate new insights [18]. Others emphasized their lack of genuine empathy, which required emotional resonance, intentional care, and moral responsibility. Although AI can mimic empathic language, it does not feel, choose to care, or assume responsibility. Without the intentionality and personal investment behind empathic acts, AI-generated responses remain fundamentally hollow [23]. Scholars adopting this view argued that until AI achieves true agency, it cannot independently deliver psychotherapy [18].
In contrast, some scholars argued that AI chatbots did not need to be treated as agents in order to demonstrate therapeutic potential [24, 25]. From this perspective, what matters is the client’s phenomenological experience of feeling supported, understood, and encouraged to reflect [25]. Research on human computer interaction shows that humans often attribute humanlike qualities to nonhuman entities through a process known as anthropomorphism, although individuals vary in this tendency [26]. Previous research found that people anthropomorphize in varying degrees and such individual differences predict the degree of connection to a nonhuman entity and the extent to which an entity is able to influence on one’s own behavior [27].Thus, the critical issue is not whether AI chatbots truly possess agency, but how users perceive and attribute agency during interactions [25]. Accordingly, empirical investigation of how individuals actually experience AI versus human support is essential to clarifying the role of agency in therapeutic change and to advancing the broader debate on AI therapists.
The need for comparative user experience research
In response to calls for more empirical research on user experiences, several recent studies have begun to examine how individuals engage with AI chatbots for mental health support. Two studies focused on users’ experiences with LLM-powered companion chatbots (e.g., Replika, Pi), but did not specify participants’ levels of mental distress [28, 29]. Another qualitative study recruited hospital outpatients and instructed them to use ChatGPT for at least 15 min per day over two weeks, but it did not report participants’ adherence to the protocol. In subsequent interviews, the study documented both positive experiences (e.g., psychoeducation) and negative ones (e.g., ethical and cultural concerns) [30]. A recent qualitative study involving participants with psychological distress explored engagement with companion chatbots like Woebot and Wysa, revealing a paradox of accessibility versus therapeutic depth [31].
Despite these early efforts, three significant limitations constrain current understanding. First, most studies rely on community samples in researcher-led settings, which may not capture naturalistic usage patterns or user-initiated interactions. Second, studies often focus exclusively on companion chatbots or conflate them with general-purpose applications, despite their distinct functional differences. Companion chatbots are a subset of AI applications specifically designed to meet users’ emotional and social needs. While often built on LLMs, they are further enhanced through interface-level features such as scripted prompts, affective cues, and customizable avatars, with the aim of simulating companionship, friendship, or even romantic connection [32, 33]. These design augmentations make it difficult to isolate the contribution of the underlying language model from the effects of additional interface elements. In contrast, studying general-purpose chatbots such as ChatGPT, which are designed for broad, open-domain interaction without domain-specific scaffolding, offers a useful starting point for exploring whether and how the core language generation capacities of LLMs might contribute to supportive dialogue in mental health contexts. Third, few studies have considered patients with experience of both AI- and human-delivered support. Patients with such dual exposure are well positioned to offer nuanced insights, as their experiences in one modality help them evaluate the perceived strengths and limitations of the other.
To address these gaps, the current study explores how individuals with mental distress experience support from both AI chatbots and human therapists, with a focus on how perceptions of agency shape these experiences. Our approach offers three advantages: First, we focused exclusively on ChatGPT as a general-purpose LLM to better understand how its core language generation capabilities might support mental health dialogue. Second, participants used ChatGPT in natural, everyday contexts, better reflecting real-world engagement. Third, all participants had dual exposure to both ChatGPT- and human-delivered support, offering a unique basis for comparative insight.
Method
Participants and recruitment
Participants were recruited using convenience and snowball sampling. Advertisements were posted on major Chinese social media platforms and online peer support groups for individuals experiencing mental distress. Interested individuals were invited to provide their contact information and complete a screening survey. The inclusion criteria were: (1) experiences mental distress, identified by either a self-reported psychiatric diagnosis or abnormal scores on any subscale of the Chinese version of the Depression Anxiety Stress Scale [34]; (2) has received psychotherapy within the past five years; (3) has used ChatGPT to address mental distress; (4) is 18 years or older; and (5) is able to speak Mandarin. The study was approved by the Human Research Ethics Committee at Hong Kong Baptist University (REC/23–24/0254). All participants were fully informed of the study objectives and provided informed consent. Participants received 50 CNY (approximately USD 7.02) after the interview as a token of appreciation.
Sixteen psychotherapy clients who met the inclusion criteria were invited to take part in an individual online interview with one of the researchers. Participants were primarily well-educated young adults (mean age = 26 years, all with a bachelor’s degree or higher) who reported receiving between 1 and 52 therapy sessions (median = 9) across various countries (i.e., China, Canada, USA) and settings (i.e., schools, hospitals, private agencies). Our sample captures a unique population at the intersection of psychotherapy delivered by a psychotherapist and AI-assisted self-help. In China, where mental health services are not publicly subsidized and are often accessed through universities or private providers, access is largely limited to individuals with higher socioeconomic status and greater mental health literacy [35]. Similarly, recent research indicates that ChatGPT users are more likely to be younger and well-educated [36]. Therefore, the participants included in this study likely reflect the early adopter profile of AI mental health chatbots in China and possibly in other similar contexts. Table 1 provides an overview of participants’ demographic information.
Participants used ChatGPT (OpenAI, San Francisco, USA), powered by the GPT-3.5 and GPT-4 language models. These models were trained on large-scale, publicly available text data and are periodically updated by OpenAI to improve performance and safety. ChatGPT incorporates built-in moderation and content-filtering mechanisms intended to block or redirect conversations involving sensitive topics such as self-harm or suicide; however, these safeguards remain limited [37, 38]. Because this study was retrospective, participants had accessed ChatGPT independently in naturalistic settings without researcher supervision, and no formal risk-management or crisis-intervention protocols were in place.
Table 1.
Participant characteristics
| Characteristics | N (%) |
|---|---|
| Age (years) | 18–36 |
| 20–29 | 11 (68.75%) |
| 30–36 | 5 (31.25%) |
| Gender | |
| Female | 10 (62.5%) |
| Male | 6 (37.5%) |
| Employment Status | |
| At school | 7 |
| Employed | 8 |
| Unemployed | 1 |
| Monthly Income (CNY) | |
| Under 5,000 | 8 (50%) |
| 5,001–10,000 | 5 (31.25%) |
| Over 10,001 | 3 (18.75%) |
| Mental Health Diagnosis a | |
| Depression | 9 (56.25%) |
| Anxiety | 3 (18.75%) |
| Bipolar | 3 (18.75%) |
| Adjustment disorder | 1 (6.25%) |
| Psychological Distress b | |
| DASS Depression | 11 (68.75%) |
| DASS Anxiety | 10 (62.5%) |
| DASS Stress | 4 (25%) |
| Currently on Medication | 5 (31.25%) |
| Therapy Delivery Location c | |
| China | 15 (88.23%) |
| Overseas | 2 (11.76%) |
| Therapy Delivery Context d | |
| Hospital | 7 (43.75%) |
| School | 7 (43.75%) |
| Private center | 8 (50%) |
| ChatGPT Version e | |
| ChatGPT 3.5 | 13 (81.25%) |
| ChatGPT 4.0 | 3 (18.75%) |
| Frequency of Usage | |
| > 3 times/week | 8 (50%) |
| 1–3 times/week | 5 (31.25%) |
| 1 time/2 week | 3 (18.75%) |
Notes: aFour participants reported more than one diagnosis; four did not report any. bindicated by scores above the normal range subscales of the Depression Anxiety and Stress Scale (DASS). cTwo participants received psychotherapy in more than one location. dSix participants received psychotherapy across multiple contexts. eFour participants used both ChatGPT 3.5 and 4.0 versions. 1 CNY = 0.14 USD
Data collection
In February 2024, XD and LLL conducted individual semi-structured interviews online in Mandarin with each participant. Each interview began with a restatement of the study’s focus and participants’ rights, followed by a semi-structured interview guide. All participants were asked a standard set of questions, while the interviewers remained flexible in using individualized probes to explore participants’ unique experiences and clarify emerging ideas. This approach allowed natural variation in dialogue while maintaining consistency in the overall interview structure.
The interview guide (see Supplementary Appendix A) was developed by XD and revised following feedback from LLL. It included questions on (1) personal mental health history (diagnosis, medication, etc.); (2) psychotherapy experience and (3) interactions with ChatGPT for addressing mental distress. When discussing interactions with human therapists and ChatGPT, participants were invited to share their motivations for seeking help, the timing and duration of their engagement, perceived changes, and any factors they believed facilitated or hindered progress. Participants who had seen multiple therapists were invited to share their experiences with their first therapist, as well as their most and least satisfactory or helpful encounters. All interviews were audio recorded using a digital recorder. Initial transcripts were generated by the recorder’s built-in transcription software and were subsequently cleaned by a research assistant and checked for accuracy by XD.
Data analysis
Interview transcripts were entered into Atlas.ti (version 24.1.1) and analyzed using reflexive thematic analysis [39, 40]. A six-phase thematic analysis process was undertaken: familiarization with the data, coding, generating initial themes, developing and reviewing themes, and refining and naming themes [41]. XD initiated the analysis by thoroughly reading the transcripts and conducting line-by-line open coding. This process was both iterative and recursive, with codes being continuously reviewed and grouped by their underlying meanings to form coherent themes and subthemes. The codes and themes were reviewed by LLL, who met with XD and engaged in reflexive discussions to interpret the data and refine the themes. Initial coding and theme development were undertaken in Chinese and then translated into English by XD and checked by LLL.
Study rigor and reflexivity
The data collection and analysis process followed the reflexive thematic analysis framework to ensure a rigorous, systematic, and reflexive analytic process [39, 40]. As Braun and Clarke highlighted, the quality of reflexive thematic analysis does not rest on consensus or reliability, but on immersion, creativity, and insight [41]. We adopted several strategies to maintain analytic quality, including ensuring sufficient information power in the dataset, allowing enough time for in-depth analysis, keeping reflexive journals, and drawing insights from supervisors and co-researchers.
A sample of 16 participants is suitable for in-depth qualitative inquiry [42]. Rather than aiming for data saturation, we followed the principle of information power [43], as reflexive TA involves open and evolving coding practices where saturation is not easily applicable [44]. According to the concept of information power, the more relevant information a sample provides for the research aim, the fewer participants are needed [43]. This study focused on a narrowly defined group to address a specific research aim—to explore how individuals with lived experience of both psychotherapy and ChatGPT use understand and compare the two modalities. To ensure the quality of conversation, all interviews were conducted by experienced researchers (XD and LLL) with relevant domain knowledge. The duration of the interviews ranged from 39 to 103 min (mean = 71.7 min), producing transcripts totaling 271,017 words (mean = 16,938.6) in length. After each interview, the interviewers wrote reflective journals documenting their own thoughts about the conversation. For example, after the 13th interview, the researcher noted, “not much new ideas, many earlier points were reaffirmed though.” Following the 15th interview, the journal recorded, “while no major new ideas were introduced, the participant shared vivid and meaningful examples.” These reflections indicated redundancy, further supporting the adequacy of information power.
Braun and Clarke also emphasize that time is a key resource for high-quality reflexive TA, highlighting the importance of the “slow wheel of interpretation” [41]. We allowed adequate time to move beyond surface-level description toward interpretive analysis. The analysis was primarily led by XD. Over a two-month period, XD immersed herself in the data by repeatedly reading transcripts and generating initial codes. In the following month, she began identifying patterned meaning across the dataset and generated candidate themes. The subsequent phase of reviewing and naming the themes took approximately five to six months and involved multiple rounds of revision and discussion. Themes and subthemes were reviewed with co-researcher LLL and qualitative expert YH, leading to restructuring of subthemes and refinement of themes. For example, the original theme one, “facilitating open sharing with lower anticipated risk” was refined to “encouraging open and authentic self-disclosure but limiting deep exploration.” In parallel, the research team engaged with relevant literature on human therapy experiences and the philosophical discussion on AI agency, which further informed the refinement and naming of themes.
In addition, the researchers acknowledged the role of subjectivity and reflexivity throughout the entire process of study design, data collection, analysis, and interpretation. Data collection and analysis primarily involved two researchers, XD and LLL, each with their personal perspectives and experiences related to the research topic. XD is a registered social worker and licensed therapist with a Ph.D. in social work. She has four years experience delivering mental health services, including psychotherapy and facilitating therapy groups. Since March 2023, she has actively used various LLM-based chatbots, including ChatGPT. LLL is a licensed therapist with a Ph.D. in social work with over ten years experience conducting psychotherapy. She uses AI chatbots occasionally in daily life. XD is optimistic about AI’s potential in psychotherapy, whereas LLL is more critical, questioning whether AI can build real relationships and facilitate therapeutic insights. This diversity of perspectives was viewed as a strength of the research team, encouraging continuous reflection on how their positionalities shaped data interpretation.
Throughout data collection and analysis, the researchers wrote memos to document emerging thoughts and notable observations and met regularly to discuss their reflections. Interpretive differences between XD and LLL were treated as opportunities for deeper engagement with the data. For instance, XD tended to stay close to participants’ language when developing themes, while LLL drew on therapeutic theory to interpret meaning. Through these discussions, the team refined and integrated both experiential and theoretical perspectives. This process enriched interpretation and reflected the principles of reflexive thematic analysis, which view researcher subjectivity as a valuable source of insight rather than a threat to validity. This study followed the MENTOR guidelines for qualitative reporting in mental health research [45], ensuring transparency and rigor in ethical inquiry, research practice, and reporting (see Supplementary Appendix B).
Results
Thematic analysis generated three themes and eight subthemes capturing what participants found useful and valuable, as well as the challenges or limitations they experienced in their interactions with ChatGPT and human therapists: (1) encouraging open and authentic self-disclosure but limiting deep exploration; (2) the myth of relationship: caring, acceptance, and understanding, and (3) fostering therapeutic change: the promise and pitfalls of data-driven solutions (see Table 2 for an overview of the findings).
Table 2.
Overview of findings
| Themes | Subthemes |
|---|---|
| 1. Encouraging open and authentic self-disclosure but limiting deep exploration | 1.1 Anonymity and consistency foster safer disclosure |
| 1.2 Autonomy as a catalyst for authentic expression | |
| 1.3 Rigid responses restrict the development of working alliance | |
| 2. The myth of relationship: caring, acceptance, and understanding | 2.1 Unconditional warmth and caring, but is it authentic? |
| 2.2 Nonjudgemental or lack of judgment? | |
| 2.3 Understanding at a literal level rather than the hermeneutic level | |
| 3. Fostering therapeutic changes: the promise and pitfalls of data-driven support | 3.1 Facilitating perspective shifts while struggling to foster intrapsychic insight |
| 3.2 Encouraging behavioral change but falling short on holistic growth |
In line with the principles of reflexive thematic analysis (Braun & Clarke, 2022), the frequency of specific experiences is represented by the number of participants who reported them. However, these numbers are descriptive and should not be interpreted as proportions of agreement, as not all participants were asked the same sub-questions; rather, the focus was on what was most relevant to their individual help-seeking experiences. Therefore, these numbers do not reflect overall agreement but instead indicate how many participants identified each experience as a meaningful aspect of their psychotherapy or interactions with ChatGPT.
Encouraging open and authentic Self-Disclosure but limiting deep exploration
Anonymity and consistency foster safer disclosure
Participants (n = 8) perceived ChatGPT as strictly bound by programming with no autonomous behavior, which fostered trust that it would adhere to privacy agreements. They reported that ChatGPT offered a heightened sense of anonymity, making them feel safer when disclosing deeply personal or potentially embarrassing issues. In contrast, eleven participants expressed concern about whether human therapists would consistently adhere to the terms of confidentiality agreements. They expressed fears of opening up, even in a supportive and confidential therapeutic environment.
Additionally, participants (n = 8) noted that ChatGPT’s consistent positive responses contributed to their sense of safety. By comparison, many described human therapists as varying widely in their styles and attitudes, which introduced uncertainty. Some participants (n = 6) raised concerns about whether the therapist was sufficiently interested or skilled to provide effective support. All participants described starting therapy with therapists as an anxiety-provoking experience, frequently recalling feelings of fear or skepticism, particularly before the first session. Participants (n = 8) mentioned being cautious about self-disclosure until trust was fully established with their therapist, with some (n = 5) indicating that fear and doubt persisted throughout the therapy process.
Autonomy as a catalyst for authentic expression
Participants (n = 10) felt empowered to lead the conversation and address their most pressing needs, as ChatGPT generates responses based on algorithms that process user prompts. They described experiencing a stronger sense of autonomy during these interactions. In contrast, participants (n = 9) often perceived an implicit or explicit power imbalance when engaging with human therapists. Four participants mentioned that their therapists sometimes dismissed their concerns, instead focusing on what the therapist believed to be more important. Others (n = 5) noted that therapists, while possessing advanced professional knowledge and skills, often employed therapeutic techniques such as asking probing questions and interpreting clients’ experiences, which inherently positioned them as authority figures. Interestingly, although ChatGPT has a vast knowledge base that far exceeds that of any individual human, its reliance on user prompts seemed to mitigate feelings of inequality that might otherwise arise from this knowledge gap. Participant #2 (female, 24) explained: “I can freely ask questions, and it [ChatGPT] will definitely not ask me any in return. It just waits for me to inquire and then responds to my queries.”
Participants also noted that the autonomy in their interactions with ChatGPT removed the pressure of managing expectations, allowing them to express themselves more authentically. A number of participants (n = 6) reported feeling pressured to make a good impression on their human therapists, leading them to self-censor during sessions. Two participants, who described themselves as socially anxious, particularly noted that they found conversations with ChatGPT to be less emotionally taxing. As Participant #14 (female, 30) stated:
When you interact with someone, for example, when I am talking to you, your mind will gradually form an impression of what kind of person I am. Even if you do not say anything, you will still form an initial impression and label me in some way. That is just part of human nature […] So, when my therapist looks at me, sometimes I feel a bit uneasy, wondering if I said something wrong. But AI does not do this.
These accounts may reflect culturally specific tensions between traditional Chinese values emphasizing deference to authority and the therapeutic ideal of client autonomy. For Chinese clients navigating these competing expectations, ChatGPT’s non-authoritative stance may have offered particular relief.
Rigid responses restrict the development of working alliance
While most participants were initially impressed by ChatGPT’s highly natural and contextually relevant responses, they noticed that conversations often became rigid as interactions continued. Four participants specifically mentioned that they quickly grew bored after recognizing its repetitive response patterns, which made it difficult to sustain deeper and more meaningful exploration.
Participants (n = 13) also criticized ChatGPT’s inability to take proactive steps when necessary, such as checking in on goals, assessing the fit of the process, or intervening when clients feel stuck or avoid key issues. For example, many participants (n = 12) described moments of uncertainty about what they were seeking, noting that ChatGPT’s reactive nature made it difficult to help them shape their reflections into goal-oriented understanding. These experiences reflected the absence of a sense of working alliance, which refers to the collaborative partnership in psychotherapy where clients and therapists negotiate goals, tasks, and shared understanding to sustain progress [46].
In human therapy, such collaboration often involves the therapist recognizing when the client feels stuck, revisiting goals, or gently guiding the process toward reflection and meaning-making. Several participants (n = 7) emphasized the value of collaborative work with human therapists to reach meaningful goals. Participant #5 (male, 36) stated: “Psychotherapy is not about right or wrong; it’s about working with a therapist to find a path that suits me.” Participant #4 (female, 31) described her experience of working with a therapist to uncover the core psychological roots of her mental health challenges:
If you asked me to speak directly, I would not know how to say it. But with his guidance […] he kept asking questions, asking me how I felt. Sometimes, as we talked about certain things, I would suddenly remember: “Oh, that happened.” […] It felt like he was helping me piece together a messy puzzle, making things clearer. I felt like I understood more about the problems I face and about myself.
The lack of collaborative negotiation in ChatGPT interactions also resembled unresolved ruptures [47] in psychotherapy. In human therapy, such moments of disconnection can deepen the working alliance when they are recognized and repaired. However, ChatGPT’s inability to detect or respond to relational strain meant that when participants felt stuck, these ruptures remained unaddressed. Consequently, conversations often stayed at a descriptive level—participants felt heard but lacked a shared sense of direction or progress that characterizes a strong working alliance.
The myth of relationship: Caring, acceptance, and understanding
Unconditional warmth and caring: is it authentic?
Participants reported that ChatGPT always responded with warmth and care, whereas some human therapists were perceived as cold or indifferent. While 11 participants described their therapists as providing comfort through warm verbal and non-verbal communication, 10 participants recounted negative experiences, feeling disrespected (n = 3), abandoned (n = 2), rejected (n = 2), or even encountering hostility (n = 2). These participants reported that their therapists lacked authentic and genuine engagement, minimized their situations and problems, or imposed their own values.
In contrast to the varied experiences with human therapists, participants noted that ChatGPT consistently offered encouragement (n = 2), reassurance (n = 6), and validation (n = 7). They used the following words to describe their interactions with ChatGPT: warm (n = 4), humane (“ren qing wei” = 4), supportive (n = 2), respectful (n = 3), patient (n = 3), and polite (n = 3). Some participants (n = 4) even felt that ChatGPT was warmer and more humane than their human therapists. According to Participant #1 stated (female, 27): “ChatGPT always validates my feelings, unlike the misunderstandings and devaluation I have experienced with my therapist. I feel that ChatGPT’s responses are more empathetic and humane, which is quite ironic.”
At the same time, participants were aware that ChatGPT’s warmth was generated through programmed responses rather than genuine emotion, which led to divided views on its authenticity. Some (n = 7) perceived ChatGPT as more genuine than human therapists, mainly because ChatGPT does not have subjective motives. They questioned the underlying motives of their therapists, with three believing that their therapists were driven by “making money,” and four feeling that their therapists “did not genuinely care”. Others (n = 3) fundamentally questioned ChatGPT’s authenticity, arguing that without “lived experiences”, it could not truly understand, validate, or value them. Participant #8 (female, 23) explained:
For me, genuine understanding requires someone to have their own thoughts and emotions […] Being understood by AI is meaningless and ineffective to me […] When human beings receive information, they think and feel something. If therapists truly understand me, they must bear the emotional costs that come with it. AI has no feelings; it merely performs tasks via machine learning without any emotional or time cost. Even if ChatGPT says, “I feel sorry for you,” who is actually feeling this sadness? Is it the hundreds of designers behind it?
Participant #8’s question—“Who is actually feeling this sadness?”—captures the central philosophical tension: therapeutic warmth derives its healing power not merely from kind words, but from the authentic presence of another consciousness bearing witness to suffering [48]. This tension may reflect a kind of cognitive dissonance, as participants experienced conflicting beliefs: “this feels supportive” versus “this entity cannot truly care.” They reconciled this tension in different ways: some redefined authenticity as freedom from ulterior motives (“it feels more genuine because it lacks selfish intent”), others accepted ChatGPT’s artificiality while valuing its functional support, and a few rejected its warmth as emotionally hollow. Overall, participants’ responses suggest that perceived authenticity is not fixed but actively constructed in the interaction. ChatGPT’s steady validation provided comfort, yet awareness of its lack of genuine feeling often tempered that comfort.
Nonjudgemental or absence of judgment?
Participants often described ChatGPT as a listener free from personal bias. Its lack of subjective values seemed to reassure them that their disclosures would not be judged. Participants (n = 13) expressed experiencing a sense of unconditional acceptance, knowing ChatGPT would never judge them. Participant #14 (female, 30) said: “Every human being has inherent biases, but ChatGPT does not, so it offers 100% acceptance.” In contrast, six participants reported feeling criticized or misunderstood because of something their therapists said or did, such as critical remarks, rushed conclusions, disdainful glances, questioning tones, and domineering postures. Participant #14 (female, 30) stated: “People have their own moral standards. When your values differ from your therapists, they might say, ‘I understand you,’ but they don’t.”
While the absence of judgment encouraged openness for some participants, it also revealed important limitations for others, particularly ChatGPT’s inability to challenge or confront users. A few participants (n = 3) criticized ChatGPT for lacking its own opinions and being merely blindly obedient. They valued human therapists’ use of professional judgment rather than simply accepting everything at face value. Participant #5 (male, 36) recounted:
My therapist, even though very professional, makes judgments about whether to interrupt me or how to proceed. But AI doesn’t do this; it just gives you what you ask for. If it [ChatGPT] agrees with everything I say, I’m not sure if this is really helping me. So, I can’t fully entrust myself to ChatGPT.
Understanding at the literal rather than the hermeneutic level
Participants (n = 13) generally felt that ChatGPT demonstrated a strong ability to understand their inputs. They reported feeling understood by ChatGPT when it voiced understanding, summarized their stories, accurately detected emotions, or expressed empathy. Despite recognizing that ChatGPT primarily reframed their inputs, participants felt it made an effort to comprehend and follow their narratives. Participant #8 (female, 23) recalled: “When I say [to ChatGPT], ‘I’m really happy today; I accomplished an important task,’ it replies, ‘Oh, I’m so happy for you…’ So, I feel like ChatGPT truly understands me and shares in my happiness.”
Some participants (n = 5) even felt that ChatGPT understood them better than their human therapists, attributing this to its access to extensive information and cultural knowledge. They believed that ChatGPT could interpret their experiences through a broader database rather than being constrained by personal experience. For instance, Participant #16 (female, 33) described her struggle with a European therapist who could not relate to her deep family ties: “He [therapist] can only understand you on a literal level because he didn’t grow up with such experiences…he’s just not in that context.” She instead felt that ChatGPT was more culturally competent: “AI sometimes doesn’t understand, but if you add context and describe it a bit, then it starts to get it; whereas with human beings, adding context is futile, they can’t learn it right away.”
However, many participants (n = 10) pointed out the limits of ChatGPT’s understanding, particularly its lack of hermeneutic understanding that involves grasping deeper meanings, emotions, and personal contexts shaping lived experience [49]. Participant #8 (female, 23) said: “I believe that ChatGPT understands the words I input, but it doesn’t truly grasp my desires or struggles as a person.” Several participants (n = 7) noted that ChatGPT struggled to understand complex or contradictory concepts and to integrate broad contextual information. Unlike human therapists, who can reflect on what lies beneath a client’s statements, ChatGPT lacks the ability to hear the unspoken, bring implicit thoughts to the surface, and explore clients’ values and worldviews to fully understand how they interpret their experiences. Furthermore, participants (n = 3) emphasized that AI chatbots have no embodied experiences and emotions, which challenges their ability to establish a connection or resonate with users through shared personal experiences.
The distinction participants drew between literal and hermeneutic understanding highlights a central question about how therapeutic insight is formed. From a hermeneutic perspective, understanding emerges through a “fusion of horizons,” in which dialogue allows both participants’ perspectives to evolve and reshape each other [49]. As Participant #5 described, therapy is like “untangling a colorful thread,” where past and present experiences are reconnected to form new meanings—a process that may depend on emotional resonance and mutual reflection beyond the current capacities of AI.
Promoting therapeutic changes: the bright and dark sides of data-driven support
Facilitating perspective shifts but struggling to foster intrapsychic insight
ChatGPT relies entirely on its database for knowledge rather than personal experience, which led participants to feel that ChatGPT provided data-driven insights free from human bias. Eight participants reported that ChatGPT offered objective explanations that helped them move beyond rigid and maladaptive thoughts and beliefs. They particularly valued ChatGPT’s epistemic superiority, as it could process data and analyses on a scale far beyond human capability. Participant #15 (male, 27) described this as “an enormous integration of human experience that allows it to more honestly reveal the essence and operational logic of matters”. One compelling example was shared by Participant #9 (female, 21), who described how ChatGPT used the metaphor of “washing clothes” to help her visualize exam anxiety:
It [ChatGPT] said that preparing for the postgraduate exam is like doing laundry in a dimly lit room. You cannot see how others are washing their clothes; you can only focus on washing your own clothes and making them as clean as possible […] I felt that ChatGPT really captured my feelings. Sometimes, I do not want to study, but I am afraid others might get ahead of me. I have no idea how far others have progressed in their preparation, and that makes me anxious […] I think this analogy captured the root of my anxiety, something I might not have been able to articulate or perhaps even realize.
Participants (n = 4) also mentioned that ChatGPT helped them adopt a more optimistic outlook, reframing their narratives by emphasizing their strengths, past successes, and available resources, shifting their focus from what is lacking or problematic to what is effective and possible. Participant #14, (female, 30) stated: “Every time after talking with ChatGPT, I felt that life isn’t as bad, people aren’t as bad, and I’m not as bad either.”
These accounts suggest that ChatGPT effectively supported cognitive and behavioral levels of change, helping users to reframe problems and generate adaptive perspectives. However, participants (n = 5) highlighted that psychotherapy was more effective than ChatGPT in facilitating deeper intrapsychic insights, which involves emotional processing, self-reflection, and the integration of past experiences. They described experiencing transformative “insight moments” during therapy, where human therapists provided professional perspectives that expanded their understanding, helped them make sense of recurring patterns within their personal histories, or encouraged critical self-reflection on previously unconscious intrapsychic conflicts. Participant #5 (male, 36) described this process as “untangling a colorful thread” and considered it to be “something ChatGPT is unable to do.”
Encouraging behavioral change but falling short on holistic growth
Nearly all participants acknowledged that the specific solutions and suggestions provided by ChatGPT were helpful. ChatGPT offers a diverse range of suggestions based on its extensive database, and participants often found at least one that suited their needs. They found its database-driven advice efficient for resolving minor psychological disturbances with minimal time, effort, or cost, helping them restore their psychological well-being. Participant #5 (male, 36) reported: “ChatGPT could quickly clear the persistent underlying noise.”
Participants also highlighted several limitations of ChatGPT compared to human therapists in facilitating long-term, meaningful personal growth. A common critique was that ChatGPT’s advice often felt overly generic (n = 11), lacking the personalized approach therapists use to create tailored action plans. Participant #7 (male, 29) remarked: “ChatGPT gave the same advice to everyone, not just me.” Despite this, participants described strategies to make their interactions with ChatGPT more useful. Ten participants mentioned asking more specific questions, such as: prompting ChatGPT to “analyze using a CBT framework” (Participant #5, male, 36). Seven others reported taking the initiative to provide more detailed context to help ChatGPT better understand their struggles. For example, Participant #15 (male, 27) described starting a new chat session specifically for psychological concerns. He shared detailed personal background information, including his key life events and psychological assessment results, and requested that ChatGPT consider these historical contexts when generating a response.
These reflections indicate that ChatGPT functioned effectively in supporting problem-solving and behavioral activation. Yet, in contrast to human therapy, its static and unidirectional format constrained the interactive process of change. Participants’ experiences suggest that ChatGPT’s support operates primarily through rational insight and data-driven reframing, without engaging users in these iterative relational processes. As a result, it seldom catalyzed the deeper restructuring of self and relationships that defines long-term therapeutic growth.
Discussion
This study is the first in-depth exploration of how individuals with mental distress experience support from both AI chatbots and human therapists, and how perceptions of agency shape these interactions. Participants generally expressed surprise at ChatGPT’s ability to communicate naturally, which far exceeded their expectations. Key findings revealed both similarities and distinctive differences between interactions with ChatGPT and human-delivered mental health support, particularly in three areas: facilitating open, authentic, and deep self-disclosure; cultivating a relationship that conveys care, acceptance, and understanding; and promoting therapeutic change. Participants’ reflections also illustrated how ambiguity around ChatGPT’s agential status shaped their interpretations of these experiences.
In the following sections, we discuss how perceived agentic qualities impact therapeutic relationships, which are widely recognized as central to therapeutic change [46, 50]. We then propose a conceptual model to illustrate how perceptions of agency bring both strengths and limitations for therapeutic engagement.
Can a therapeutic relationship be established without human agency?
Our findings suggest that many participants perceived their interactions with ChatGPT as similar to therapy sessions in several important ways. A significant observation was that participants reported feeling understood, accepted, and cared for by ChatGPT, which reflects key aspects of a therapeutic relationship [51]. This aligns with prior work showing that AI-generated responses in emotionally complex scenarios can make recipients feel more understood than responses from untrained humans when AI’s identity is concealed [52]. In our study, some participants reported that even when they knew they were interacting with AI, ChatGPT was at times perceived as providing more supportive responses than trained professionals.
Nonetheless, we could not ascertain that AI chatbots are capable of building therapeutic relationships in the same way as human therapists. Participants frequently described a subtle sense of detachment linked to ChatGPT’s artificial nature. Even those who felt “upheld and touched” or that ChatGPT had “walked into my heart” also reported moments of loss when reminded it was governed by algorithms. This ambiguity reflects an underlying tension that may stem from the hybrid status of AI chatbots—artifacts that exhibit both tool-like and agent-like features yet do not meet the full criteria for agency [18]. As some scholars have noted, while AI can accurately identify users’ emotions (i.e., cognitive empathy), it cannot share these emotions (i.e., emotional empathy) or demonstrate genuine concern for others (i.e., motivational empathy) [23]. These inherent tensions resonate with theoretical that question whether the therapeutic relationship can be deconstructed into discrete elements, such as warmth, empathy, and acceptance, and whether these elements would retain their effectiveness if delivered by a nonhuman entity, such as AI [53]. While these questions are beyond the scope of this study, our findings highlight the importance of users’ perceptions of agency in shaping how they engage with AI chatbots for mental health support.
Human versus AI support: the double-edged sword of agential status
Despite the heterogeneity of participants’ experiences interacting with human therapists and ChatGPT to address mental distress, a common thread emerged; participants valued having a safe space to freely share their concerns with an entity that they felt understood, accepted, and cared for them, helping them make sense of their difficulties and guiding them toward achieving meaningful therapeutic change. Participants’ accounts revealed that ChatGPT addressed their mental health needs in ways distinct from human therapists. These differences were closely tied to its non-agential features, acting as a double-edged sword, bringing both benefits and challenges to users’ experiences (as illustrated in Fig. 1).
Fig. 1.
The Double-edged sword of agential features
First, participants perceived that ChatGPT demonstrated responsive behaviors rather than autonomous, self-directed behaviors. This non-agential nature gave users more control over interactions and made them feel safe. Participants reported that ChatGPT provided consistent positive responses and believed it could maintain confidentiality. However, the dependence on users’ prompts meant that ChatGPT cannot initiate inquiries, provide guidance when participants felt stuck, or adapt therapeutic strategies to individual needs, which significantly limits the communication efficiency, exploration depth, and problem-solving effectiveness.
Second, participants noted an absence of intentions and motives in ChatGPT’s interactions. Therefore, they viewed ChatGPT as providing help with sincerity and selflessness, free from the personal biases or self-interest they sometimes perceived in human therapists. However, this lack of intentionality also meant there was an absence of “goodwill,” leading some participants to doubt the genuineness of ChatGPT’s responses.
Third, participants perceived that ChatGPT maintained an objective stance rather than holding subjective values or opinions. ChatGPT’s perceived objectivity provided unconditional acceptance and objective suggestions free from personal biases. Nonetheless, this all-accepting nature may hinder its ability to offer critical judgment when necessary, making it less effective in guiding clients through complex challenges requiring reality-testing or confrontation.
Fourth, participants described ChatGPT as operating from a disembodied state, drawing on extensive databases rather than relying on embodied experiences or feelings. Participants acknowledged ChatGPT’s cognitive strengths in providing data-driven insights across diverse mental health challenges and cultural contexts. However, its disembodied state means it could not truly relate to lived experiences or provide empathy grounded in shared human emotions or experiences.
In contrast, participants viewed human therapists as agents whose autonomy and subjectivity introduced both unpredictability and greater potential in the therapeutic process. On one hand, some participants reported difficulties: struggles in establishing a trusting relationship with therapists, feelings of rejection rather than understanding, and even perceived therapy as more harmful than helpful. On the other hand, these intrinsic human attributes equip therapists to engage in profound and meaningful therapeutic work. Skilled therapists can transform initial fears into safety, cultivating relationships that allow clients to explore pathways for meaningful change and growth. This intentional relationship-building carries intrinsic healing value [54], that may surpass algorithm-driven responses from AI.
Clinical implications: human agency and AI non-agency as complementary strengths
The findings empirically extend the ongoing debate on whether AI chatbots must truly possess agency to facilitate therapeutic change. Participants’ experiences cannot be captured by a simple dichotomy of possessing or lacking agency. Instead, they reveal a dynamic process in which users continuously negotiated ChatGPT’s humanlike and mechanical qualities to make sense of their interactions. At times, participants attributed humanlike traits such as emotion and intention to ChatGPT, which provided comfort and encouragement. Yet they also remained aware of its artificial nature, which fostered a distinct sense of safety, predictability, and control. This perspective reframes the discussion from whether AI can be a therapist to how the interplay between agential and non-agential features generates new forms of therapeutic meaning.
Given the double-edged nature of agency, achieving the right balance between humanlike design and non-agential features may be crucial for enhancing the effectiveness of digital mental-health interventions. Participants’ attributions of emotion and intention to ChatGPT may reflect partial activation of theory-of-mind processes, leading to varying degrees of anthropomorphic experience. Previous research found that individuals differ in their tendency to anthropomorphize, and such differences predict their emotional connection to nonhuman entities and the degree to which those entities influence their behavior [26, 27]. Research on intelligent technologies such as voice assistants and autonomous vehicles further indicates that the more humanlike a system appears, the more people trust it to perform competently, regardless of the moral valence of its task [55, 56]. It is therefore unsurprising that many mental-health chatbots are deliberately designed with humanlike cues to evoke empathy and engagement [57]. However, our findings indicate that non-agential qualities—neutrality, predictability, and lack of judgment—were equally valued in therapeutic contexts. Designs that overemphasize anthropomorphic traits to enhance perceived agency may blur these advantages and produce unintended effects. For instance, AI chatbots with pronounced personalities may undermine patients’ perception of objectivity, whereas overly proactive chatbots might diminish users’ sense of autonomy during interactions. Anthropomorphic representations should therefore be incorporated with caution in the design of mental-health chatbot, preserving the benefits of non-agency while maintaining user engagement and trust.
The double-edged nature of agency also resonated with challenges in human-delivered psychotherapy, inviting reflection on the ways in which therapists’ agency can both enable and hinder change. Consistent with systematic reviews and meta-analyses synthesizing global evidence on negative experiences in therapy [58, 59], our participants described a range of similar difficulties in their therapy sessions, alongside helpful and insightful moments. Many negative experiences stemmed from the double-edged nature of therapists’ agential features (as discussed above). While clinicians should leverage the strengths of their agentic qualities, they must also remain mindful of associated risks, such as attending to power dynamics and ensuring clients feel empowered throughout the therapeutic process.
Taken together, these findings point to the need for an integrated approach that draws on the complementary strengths of human agency and AI non-agency. Rather than positioning AI as a competitor, mental health systems might consider how its beneficial features can complement therapeutic services. Based on our findings, we recommend two integration strategies. First, AI chatbots could be deployed as preliminary screening and psychoeducation tools. Their consistent, non-judgmental tone encourages self-disclosure and may reduce anxiety during early encounters. Clinicians might incorporate chatbots into intake processes to help clients organize and articulate their concerns before sessions. Second, AI chatbots may serve as between-session support tools under human oversight, offering immediate assistance through emotion regulation strategies, cognitive reframing prompts, or structured exercises. All participants appreciated ChatGPT’s ability to deliver timely and diverse coping strategies. Some also noted its capacity to provide structured guidance when specifically prompted (e.g., “analyze using a CBT framework”), underscoring its potential for guiding structured exercises.
Ethical considerations and risk management
Using AI chatbots as therapists comes with concerning risks, particularly in unguided use scenarios. First, ChatGPT’s “blind obedience” poses risks for vulnerable populations. While this responsiveness felt accepting to participants with intact reality testing, it also points to a risk identified when interacting with individuals experiencing cognitive distortions or delusional thinking, such uncritical agreement may inadvertently reinforce distorted beliefs [60]. Recent research reported that AI chatbots can respond inappropriately to mental health crises, potentially reinforcing suicidal ideation [61]. Unlike human therapists, who can recognize and gently challenge maladaptive thoughts, ChatGPT’s consistently affirmative stance may validate delusions or maladaptive behaviors. This risk may be heightened among individuals with ADHD, OCD, or substance use disorders, where uncritical agreement could strengthen avoidance patterns or impulsive tendencies [7].
Second, the autonomy participants valued creates oversight gaps when users cannot recognize their limitations or crisis situations —a challenge particularly common among individuals experiencing mental distress [60]. Participants noted ChatGPT’s inability to “check in on goals” or “intervene when clients feel stuck”. For individuals in acute distress or those with impaired judgment, this absence of proactive monitoring constitutes a safety risk. As Beg [62] emphasized in the guidelines for responsible AI integration in mental health, oversight mechanisms and clear escalation protocols are essential but often absent in publicly available chatbots.
Third, emotional dependency and therapeutic misconceptions may pose longer-term risks. Participants described ChatGPT’s “unconditional warmth” and constant availability as comforting, yet these features may foster overreliance or parasocial attachment. Such relationships could substitute for genuine human connection, exacerbate social isolation [61, 62], reduce conflict-resolution capacity [64], or discourage help-seeking from professionals [63]. These risks are especially salient for vulnerable groups such as minors and individuals with limited digital literacy, who may overestimate AI’s therapeutic capacity and develop unrealistic expectations [64].
Fourth, individuals with certain personality vulnerabilities, particularly Borderline Personality Disorder, may face increased risks. These individuals may be more susceptible to forming maladaptive relational patterns with an AI that offers steady warmth and affirmation without the ruptures and repairs that occur in real relationships. This may reinforce idealization, undermine distress tolerance, and interfere with developing stable interpersonal relationships. Clinicians should be careful when considering the use of AI chatbots for individuals with personality pathology.
Despite increasing calls for clear guidelines on the use of AI chatbots in mental health [65], many AI mental health chatbots remain publicly accessible with minimal oversight, interacting with millions of users worldwide. Just as most countries require licensed mental health professionals to undergo rigorous training, certification, and supervision, mental health chatbots should also be subject to similar evaluation and regulatory processes. We recommend establishing an interdisciplinary evaluation framework that includes regular safety assessments, transparency about system limitations, and clear protocols for crisis escalation and professional referral. The clinical potential identified in this study can only be realized when risks are explicitly acknowledged and addressed through responsible design, proactive monitoring, and appropriate regulation. Without these protections, the very features that make AI chatbots accessible and engaging may also make them unsafe.
Notably, this study also highlights ethical challenges concerning accessibility and equity in AI-assisted mental health care. The sample consisted mainly of well-educated young Chinese participants with relatively high levels of digital and mental health literacy. Many played an active role in co-creating therapeutic value by adapting prompts, refining questions, and drawing on their prior knowledge of psychotherapy to steer conversations toward personally meaningful directions. This dual literacy may partly explain the positive experiences reported, as participants were able to elicit more tailored and supportive responses.
Although digital mental health solutions, including AI chatbots, are often promoted as tools to bridge gaps in low-resource settings [66], our findings indicated that variations in users’ digital competence and therapeutic experience may influence how effectively they engage with these systems. Individuals with lower literacy or fewer educational opportunities may therefore face structural disadvantages, potentially reinforcing existing disparities in access to care. Future research should explore how design improvements such as adaptive interfaces, guided prompting systems, and culturally sensitive dialogue models can promote more equitable engagement across diverse user groups. In parallel, public education initiatives to strengthen both AI literacy and mental health literacy are essential to ensure that underserved populations can benefit from mental health chatbots.
Limitations and future studies
Our findings provide valuable insights into the potential role of AI chatbots in mental health support, but they should be interpreted with caution. First, clients’ expectations play a significant role in psychotherapy [46]. The high satisfaction with ChatGPT reported in this study may partly reflect participants’ lower expectations of AI chatbots compared with trained professionals. We did not systematically assess participants’ therapeutic goals. In China, limited mental health literacy means many individuals are unfamiliar with different counseling approaches and tend to hold problem-focused expectations [35]. Yet participants varied considerably in what they sought—from “clearing persistent underlying noise” (Participant #36) to “untangling a colorful thread” (Participant #5)—indicating that individuals approached ChatGPT with differing expectations and therapeutic aims. Future research should examine how therapeutic goals moderate AI chatbot effectiveness.
Second, our sample consisted of well-educated, young Chinese participants with relatively high digital and mental health literacy, which significantly limits transferability. These participants actively co-created therapeutic value by providing detailed context, refining prompts, and guiding interactions—a form of “therapeutic prompt literacy” supported by their prior education and therapy experience. Individuals with lower literacy, no therapy experience, or cognitive impairment may find it difficult to engage in this way, raising equity concerns. While our sample represents early adopters in China, it remains unclear whether AI chatbots can effectively serve users with limited digital or therapeutic literacy, an important issue given that such populations account for most unmet mental health needs. While this group likely represents the primary demographic currently using AI mental health chatbots in China and similar contexts, findings may not extend to individuals from different cultural backgrounds, education levels, or levels of technology literacy. This profile therefore represents both a strength, in capturing early adopters, and a limitation for broader applicability.
An additional limitation concerns the heterogeneity of participants’ prior psychotherapy experiences. Participants engaged in different therapeutic approaches, showed wide differences in treatment duration, and pursued goals that ranged from symptom relief to deep intrapsychic exploration. This heterogeneity constitutes a confounding variable when comparing experiences with human therapists versus ChatGPT.
Third, although the sample size was appropriate for qualitative inquiry [42] and achieved strong information power [43], it limits our ability to explore subgroup differences. Such differences may be particularly relevant given the influence of individual client characteristics on therapeutic outcomes [46]. Moreover, the qualitative design limited our ability to examine how the key elements in the conceptual model interact dynamically. Future studies could build on this work by investigating how AI systems process and respond to users’ psychological states. Applying theoretical perspectives such as theory of mind may further clarify how these mechanisms influence the depth and quality of AI–human therapeutic interactions.
Fourth, cultural context strongly influenced participants’ experiences and interpretations, which limits the transferability of findings beyond Chinese or similar collectivist settings. The importance participants attached to ChatGPT’s non-judgmental stance, anonymity, and reduced power asymmetry likely reflects culturally specific concerns about stigma and hierarchical relationships. As mental health literacy, help-seeking norms, and views on autonomy and acceptance differ across cultures, the advantages and limitations of AI chatbots over human therapy identified here may not apply to individualistic societies. Cross-cultural research is needed to clarify AI’s role in diverse therapeutic contexts.
Additionally, the retrospective design comparing AI interactions with therapy experiences introduces potential recall bias and temporal confounding. Participants may have used ChatGPT and therapy for different concerns, at different levels of severity, and in different life contexts, making direct comparisons challenging. Finally, this study focused exclusively on ChatGPT as a general-purpose LLM and may have limited transferability to other AI chatbots with different design features or capabilities.
Conclusions
Our findings indicate that AI chatbots and human therapists support individuals with mental distress in distinct ways, largely reflecting their fundamental differences in agential status. This study also reveals the double-edged nature of agential status in psychotherapy, challenging the assumption that therapeutic change inherently requires communication between two conscious agents. Rather than attempting to make AI chatbots more human-like, future mental health services may benefit from integrated care models that harness the non-agentic strengths of AI alongside the agentic qualities of human therapists. Further cross-cultural and cross-population research is needed to validate and extend these findings in diverse contexts.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
We are deeply grateful to all participants who generously shared their experiences. Their trust and openness made this study possible. We also thank Dr. J. P. Grodniewicz (Institute of Philosophy, Jagiellonian University) and Dr. Lik-Hang Lee (Department of Computing, The Hong Kong Polytechnic University) for their insightful feedback, which greatly improved the manuscript.
Author contributions
XD designed the study, collected and analyzed the data, and drafted the manuscript; LLL contributed to the study conceptualization, collected data, reviewed the themes, and write the manuscript; YL collected and curated the data; Y-TH consulted on the themes and edited the manuscript; DFKW consulted on the themes and edited the manuscript. All authors approved the final version of the manuscript.
Funding
This research did not receive any grant from funding agencies in the public, commercial, or not-for-profit sectors.
Data availability
Deidentified data are available upon request.
Declarations
Ethics approval and consent to participate
This study was approved by the Human Research Ethics Committee of Hong Kong Baptist University (REC/23–24/0254). Written and oral informed consent was obtained from all participants.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Vigo D, Jones L, Atun R, Thornicroft G. The true global disease burden of mental illness: still elusive. Lancet Psychiatry. 2022;9(2):98–100. [DOI] [PubMed] [Google Scholar]
- 2.Auerbach RP, Mortier P, Bruffaerts R, Alonso J, Benjet C, Cuijpers P, Demyttenaere K, Ebert DD, Green JG, Hasking P et al. WHO World Mental Health Surveys International College Student Project: Prevalence and Distribution of Mental Disorders. Journal of abnormal psychology (1965) 2018, 127(7):623–638. [DOI] [PMC free article] [PubMed]
- 3.Evans-Lacko S, Aguilar-Gaxiola S, Al-Hamzawi A, Alonso J, Benjet C, Bruffaerts R, Chiu WT, Florescu S, de Girolamo G, Gureje O, et al. Socio-economic variations in the mental health treatment gap for people with anxiety, mood, and substance use disorders: results from the WHO world mental health (WMH) surveys. Psychol Med. 2018;48(9):1560–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Andrade LH, Alonso J, Mneimneh Z, Wells JE, Al-Hamzawi A, Borges G, Bromet E, Bruffaerts R, de Girolamo G, de Graaf R, et al. Barriers to mental health treatment: results from the WHO world mental health surveys. Psychol Med. 2014;44(6):1303–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Radziwill NM, Benton MC. Evaluating quality of chatbots and intelligent conversational agents. In. Ithaca: Cornell University Library, arXiv.org;; 2017. [Google Scholar]
- 6.Beg MJ, Verma M, Verma MVCKM. Artificial intelligence for psychotherapy: A review of the current state and future directions. Indian J Psychol Med. 2025;47(4):314–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Beg MJ, Verma MK. Exploring the potential and challenges of digital and AI-Driven psychotherapy for ADHD, OCD, Schizophrenia, and substance use disorders: A comprehensive narrative review. Indian J Psychol Med 2024:02537176241300569. [DOI] [PMC free article] [PubMed]
- 8.Li H, Zhang R, Lee Y-C, Kraut RE, Mohr DC. Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being. Npj Digit Med. 2023;6(1):236. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Thorsten Brants ACP, Xu P, Och FJ, Dean J. Large Language Models in Machine Translation. In: Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning: 2007; 2007: 858–867.
- 10.Vaswani A. Attention is all you need. Adv Neural Inf Process Syst 2017.
- 11.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I. Attention is all you need. In. Ithaca: Cornell University Library, arXiv.org;; 2023. [Google Scholar]
- 12.Ray PP. ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet Things Cyber-Physical Syst. 2023;3:121–54. [Google Scholar]
- 13.Sufyan NS, Fadhel FH, Alkhathami SS, Mukhadi JY. Artificial intelligence and social intelligence: preliminary comparison study between AI models and psychologists. Front Psychol. 2024;15:1353022. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, Faix DJ, Goodman AM, Longhurst CA, Hogarth M. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Vowels LM. Are chatbots the new relationship experts? Insights from three studies. Computers Hum Behavior: Artif Hum 2024:100077.
- 16.Galido PV, Butala S, Chakerian M, Agustines D. A case study demonstrating applications of ChatGPT in the clinical management of treatment-resistant schizophrenia. Cureus 2023, 15(4). [DOI] [PMC free article] [PubMed]
- 17.Grodniewicz J, Hohol M. Therapeutic chatbots as cognitive-affective artifacts. Topoi 2024:1–13.
- 18.Sedlakova J, Trachsel M. Conversational artificial intelligence in psychotherapy: a new therapeutic tool or agent? Am J Bioeth. 2023;23(5):4–13. [DOI] [PubMed] [Google Scholar]
- 19.Ladmanová M, Řiháček T, Timulak L, Jonášová K, Kubantová B, Mikoška P, Polakovská L, Elliott R. Client-identified outcomes of individual psychotherapy: a qualitative meta-analysis. Lancet Psychiatry. 2025;12(1):18–31. [DOI] [PubMed] [Google Scholar]
- 20.Markus S. Agency. In: The Stanford Encyclopedia of Philosophy. Edited by Zalta EN: The Metaphysics Research Lab, Center for the Study of Language and Information, Stanford University, Stanford, CA 94305 – 4115; 2019.
- 21.Sapkota R, Roumeliotis K, Karkee M. AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges. arXiv 2025. arXiv preprint arXiv:250510468.
- 22.Martela F. Artificial intelligence and free will: generative agents utilizing large Language models have functional free will. Ai Ethics (Online). 2025;5(4):4389–400. [Google Scholar]
- 23.Perry A. AI will never convey the essence of human empathy. Nat Hum Behav. 2023;7(11):1808–9. [DOI] [PubMed] [Google Scholar]
- 24.Grodniewicz JP, Hohol M. Therapeutic conversational artificial intelligence and the acquisition of Self-understanding. Am J Bioeth. 2023;23(5):59–61. [DOI] [PubMed] [Google Scholar]
- 25.Hurley ME, LB H, Smith JN. Therapeutic artificial intelligence: does agential status matter? Am J Bioeth. 2023;23(5):33–5. [DOI] [PubMed] [Google Scholar]
- 26.Epley N, Waytz A, Cacioppo JT. On seeing human: A three-factor theory of anthropomorphism. Psychol Rev. 2007;114(4):864–86. [DOI] [PubMed] [Google Scholar]
- 27.Waytz A, Cacioppo J, Epley N. Who sees human? The stability and importance of individual differences in anthropomorphism. Perspect Psychol Sci. 2010;5(3):219–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Maples B, Cerit M, Vishwanath A, Pea R. Loneliness and suicide mitigation for students using GPT3-enabled chatbots. Npj Mental Health Res. 2024;3(1):4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Siddals S, Torous J, Coxon A. It happened to be the perfect thing: experiences of generative AI chatbots for mental health. Npj Mental Health Res. 2024;3(1):48. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Alanezi F. Assessing the effectiveness of ChatGPT in delivering mental health support: a qualitative study. J Multidisciplinary Healthc 2024:461–71. [DOI] [PMC free article] [PubMed]
- 31.Beg MJ, Verma MK. Artificial Intelligence-based psychotherapy: a qualitative exploration of usability, Personalization, and the perception of therapeutic progress. Indian J Psychol Med. 2025:02537176251357477. [DOI] [PMC free article] [PubMed]
- 32.Malfacini K. The impacts of companion AI on human relationships: risks, benefits, and design considerations. Ai & Society; 2025.
- 33.Skjuve M, Følstad A, Fostervold KI, Brandtzaeg PB. My chatbot Companion - a study of Human-Chatbot relationships. Int J Hum Comput Stud. 2021;149:102601. [Google Scholar]
- 34.Wang K, Shi H-S, Geng F-L, Zou L-Q, Tan S-P, Wang Y, Neumann DL, Shum DHK, Chan RCK. Cross-cultural validation of the Depression Anxiety Stress Scale–21 in China. In., vol. 28. US: American Psychological Association; 2016: e88-e100. [DOI] [PubMed]
- 35.Ding R, He P, Zheng X. Socioeconomic inequality in rehabilitation service utilization for schizophrenia in china: findings from a 7-year nationwide longitudinal study. Front Psychiatry. 2022;13:914245. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Kacperski C, Ulloa R, Bonnay D, Kulshrestha J, Selb P, Spitz A. Characteristics of ChatGPT users from germany: implications for the digital divide from web tracking data. PLoS ONE. 2025;20(1):e0309047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.GPT-4 System Card. [https://cdn.openai.com/papers/gpt-4-system-card.pdf]
- 38.Usage Policies. [ https://openai.com/zh-Hans-CN/policies/usage-policies]
- 39.Braun V, Clarke V. Reflecting on reflexive thematic analysis. Qualitative Res Sport Exerc Health. 2019;11(4):589–97. [Google Scholar]
- 40.Braun V, Clarke V. One size fits all? What counts as quality practice in (reflexive) thematic analysis? Qualitative Res Psychol. 2021;18(3):328–52. [Google Scholar]
- 41.Braun V, Clarke V. Thematic analysis: a practical guide. Los Angeles: SAGE; 2022. [Google Scholar]
- 42.Guest G, Bunce A, Johnson L. How many interviews are enough? An experiment with data saturation and variability. Field Methods. 2006;18(1):59–82. [Google Scholar]
- 43.Malterud K, Siersma VD, Guassora AD. Sample size in qualitative interview studies: guided by information power. Qual Health Res. 2016;26(13):1753–60. [DOI] [PubMed] [Google Scholar]
- 44.Braun V, Clarke V. To saturate or not to saturate? Questioning data saturation as a useful concept for thematic analysis and sample-size rationales. Qualitative Res Sport Exerc Health. 2021;13(2):201–16. [Google Scholar]
- 45.Beg MJ. Qualitative methods in mental health research: standards for ethical Inquiry, research Practice, and peer review. Indian J Psychol Med 2025:02537176251363817. [DOI] [PMC free article] [PubMed]
- 46.Wampold BE, Imel ZE. The great psychotherapy debate: the evidence for what makes psychotherapy work. Routledge; 2015.
- 47.O’Keeffe S, Martin P, Midgley N. When adolescents stop psychological therapy: Rupture–repair in the therapeutic alliance and association with therapy ending. Psychotherapy. 2020;57(4):471. [DOI] [PubMed] [Google Scholar]
- 48.Greenberg LS, Geller S. Congruence and therapeutic presence. Rogers’ Therapeutic Conditions: Evol Theory Pract. 2001;1:131–49. [Google Scholar]
- 49.Clark J. Philosophy, Understanding and the consultation: a fusion of horizons. Br J Gen Pract. 2008;58(546):58. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Castonguay LG, Hill CE. Insight in psychotherapy. 1st ed. Washington, DC: American Psychological Association; 2007. [Google Scholar]
- 51.Beck AT, Rush AJ, Shaw BF, Emery G, DeRubeis RJ, Hollon SD. Cognitive therapy of depression. Guilford; 2024.
- 52.Yin Y, Jia N, Wakslak CJ. AI can help people feel heard, but an AI label diminishes this impact. Proc Natl Acad Sci. 2024;121(14):e2319112121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Grodniewicz J, Hohol M. Waiting for a digital therapist: three challenges on the path to psychotherapy delivered by artificial intelligence. Front Psychiatry. 2023;14:1190084. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Hartford A, Stein DJ. The machine speaks: conversational AI and the importance of effort to relationships of meaning. JMIR Ment Health. 2024;11:e53203. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Waytz A, Heafner J, Epley N. The Mind in the machine: anthropomorphism increases trust in an autonomous vehicle. J Exp Soc Psychol. 2014;52:113–7. [Google Scholar]
- 56.Ruijten PAM, Terken JMB, Chandramouli SN. Enhancing trust in autonomous vehicles through intelligent user interfaces that mimic human behavior. Multimodal Technol Interact. 2018;2(4):62. [Google Scholar]
- 57.Monteith S, Glenn T, Geddes JR, Whybrow PC, Achtyes E, Bauer M. Anthropomorphic technology in everyday life: focus on chatbots and impacts on mental health. Eur Arch Psychiatry Clin NeuroSci 2025. [DOI] [PMC free article] [PubMed]
- 58.Li ACM, Mak WWS. Service users’ perspective of therapist-related unwanted events in psychotherapy—A systematic review. J Couns Psychol. 2025;72(4):390–401. [DOI] [PubMed] [Google Scholar]
- 59.Vybíral Z, OB M, Tomáš Ř, Barbora U, Gocieková V. Negative experiences in psychotherapy from clients’ perspective: A qualitative meta-analysis. Psychother Res. 2024;34(3):279–92. [DOI] [PubMed] [Google Scholar]
- 60.Hansen CF, Torgalsbøen A-K, Røssberg JI, Andreassen OA, Bell MD, Melle I. Object relations and reality testing in schizophrenia, bipolar disorders, and healthy controls: differences in profiles and clinical correlates. Compr Psychiatr. 2012;53(8):1200–7. [DOI] [PubMed] [Google Scholar]
- 61.Moore JG. Declan and Agnew, William and Klyman, Kevin and Chancellor, Stevie and Ong, Desmond C. and Haber, Nick: Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. In: ACM Conference on Fairness, Accountability, and Transparency: 2025: ACM; 2025.
- 62.Beg MJ. Responsible AI integration in mental health research: Issues, Guidelines, and best practices. Volume 47. New Delhi, India: SAGE Publications Sage India; 2025. pp. 5–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Denecke K, Abd-Alrazaq A, Househ M. Artificial intelligence for chatbots in mental health: opportunities and challenges. Multiple Perspect Artif Intell Healthcare: Opportunities Challenges 2021:115–28.
- 64.Solyst J, Yang E, Xie S, Hammer J, Ogan A, Eslami M. Children’s overtrust and shifting perspectives of generative AI. 905–12. In.; 2024.
- 65.De Freitas J, Cohen IG. The health risks of generative AI-based wellness apps. Nat Med. 2024;30(5):1269–75. [DOI] [PubMed] [Google Scholar]
- 66.Wang X, Sanders HM, Liu Y, Seang K, Tran BX, Atanasov AG, Qiu Y, Tang S, Car J, Wang YX et al. ChatGPT: promise and challenges for deployment in low- and middle-income countries. Lancet Reg Health – Western Pac 2023, 41. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Deidentified data are available upon request.

