Abstract
Autonomy is a fundamental ethical principle in artificial intelligence (AI) ethics. Current discussions regarding autonomy-related risks in human–AI interaction, as well as potential mitigation strategies, have mainly focused on recommendation systems and algorithmic decision-making systems. However, systematic analyses of the autonomy issues posed by newly emerging generative AI or chatbots (such as ChatGPT) remain scarce. This paper aims to bridge this gap and proposes a Socratic method-based chatbot, herein designated SocrAI, as a potential countermeasure informed by recent technological advancements. We identify two primary forms of autonomy risk associated with generative chatbots—false mental states and cognitive deskilling—and, through an examination of their underlying causes, argue that the Socratic method offers a plausible means of mitigation. The paper further assesses the feasibility of employing the Socratic method within generative chatbots to preserve and support users’ autonomy and outlines the prospective implementation of SocrAI together with directions for future work. SocrAI represents a novel attempt to strengthen human agency, one that encourages users to become self-initiating inquirers who, with the assistance of AI, actively engage their abilities in the pursuit of answers.
Keywords: Generative AI, Chatbot, Autonomy, The Socratic method, Human agency
Introduction
Autonomy, as a distinct human objective of intrinsic value, is intricately tied to human dignity, responsibility, and self-realization. It not only shapes the developmental trajectory of individual personhood but also constitutes a fundamental end of human existence. Given its significant ethical implications, autonomy has been recognized as an integral of artificial intelligence (AI) ethics and a foundational principle in the design and use of human-centered AI (Prunkl, 2024; Jobin et al., 2019; Floridi & Cowls, 2019). Nonetheless, the interactions between humans and AI systems continue to present numerous challenges to user autonomy, including the unauthorized collection of personal data, online manipulation and behavioral nudging, as well as concerns regarding algorithmic opacity and bias (Botes, 2023; Héder, 2023; Fazelpour & Danks, 2021). These challenges may result in a misalignment between users’ goals and their actions, and in certain instances, may impair their cognitive and critical thinking abilities (Prunkl, 2024; Lu, 2024; Bonicalzi et al., 2023).
Current discussions regarding the autonomy risks in AI and potential countermeasures have predominantly focused on algorithmic recommendation systems—such as YouTube, Twitter, TikTok, and Amazon—that provide content and product suggestions based on user preferences (Tiribelli & Calvaresi, 2024; Laitinen & Sahlgren, 2021; Del Valle & Lara, 2024), as well as on algorithmic decision-making systems that assist human decision processes in domains including medical diagnosis, autonomous driving, judicial assessment, and commerce (Steyvers & Kumar, 2024; Fossa, 2024; Karlan, 2024). The emergence of ChatGPT has not only expedited technological advancements in AI but has also fundamentally altered public perceptions of AI and the interaction patterns with it. Generative AI models or chatbots, such as ChatGPT1 and Gemini, exhibit enhanced task generalizability and provide more natural, human-like interactions, thereby enriching the personalized experience for users. They are adept at executing a wide range of tasks, including translation, creative writing, question answering, and programming, while assuming various conversational roles. However, as an emergent AI system, ChatGPT inevitably raises various ethical concerns in its interactions with users, encompassing issues of privacy, plagiarism and copyright attribution, misinformation, and malicious use (Weidinger et al., 2022; Gabriel et al., 2024; Zhou et al., 2024; Hua et al., 2024). While scholars have initiated investigations into these issues, the autonomy-related risks associated with generative AI systems remain systematically underexamined.
Ensuring alignment between AI and human values has consistently been a central focus within the field of AI ethics. In response to the (ethical) challenges posed by (generative) AI systems, policymakers, ethicists, and technical experts have proposed a diverse array of constructive strategies. From a governance and regulatory standpoint, efforts have focused on the establishment of comprehensive legal and institutional frameworks that explicitly delineate the responsibilities, rights, and obligations of both developers and users. These frameworks are designed to facilitate the lawful and ethical use of generative AI, particularly concerning intellectual property rights, data security, and privacy protection (Hacker et al., 2023; Dixon, 2023). Other proposals advocate for the creation of oversight mechanisms and dedicated review bodies to monitor the development and application of LLMs (Mökander et al., 2024). Additionally, businesses, social organizations, and the general public are urged to participate in collaborative regulatory processes (Stahl et al., 2022). On the ethical front, scholars and practitioners have underscored the significance of transparency and explainability in AI-generated content, the mitigation of bias and discrimination, and the protection of individual rights. Complementary principles—such as inclusivity and sustainability—are also deemed essential for the development of responsible, trustworthy, and human-centered AI systems (Harrer, 2023; Oniani et al., 2023). At the individual level, enhancing algorithmic literacy and awareness is imperative. Education and training initiatives are encouraged to equip users with knowledge necessary to comprehend the mechanisms underlying algorithmic outputs, identify potential biases or inaccuracies in AI, and anticipate associated ethical risks (Ienca, 2023; Höller et al., 2023; Dogruel, 2021). Armed with this knowledge, individuals are better positioned to manage and protect their personal data, critically engage with AI-generated content and decision-support tools, verify the reliability of information, and ultimately retain control over their decision-making processes (Zhou et al., 2024; Steinerová, 2023).
While these measures are undeniably effective in mitigating certain challenges posed by generative AI, structural factors continue to undermine user autonomy. These factors do not stem from deliberate design choices by developers; rather, they arise from the intrinsic limitations of AI systems—particularly in the mechanisms of data collection and probabilistic content generation—as well as the human inclination, as boundedly rational agents, to pursue enhanced capacities and reduced cognitive burdens (see Nyholm, 2024). Despite the aspiration for AI to learn and think like humans, AI ultimately lacks consciousness and intentional thought, generating outputs based on statistical inference rather than genuine understanding. This fundamental discrepancy creates a persistent tension, wherein user autonomy is continually at risk of being compromised. Under prevailing human–AI interaction models, the erosion of autonomy appears to be an inevitable cost of maximizing cognitive efficiency and performance (Krook, 2025; Schaap et al., 2024). More broadly, the dominant technocentric paradigm in AI development tends to regard increased capability as a panacea for the ethical issues posed by AI. While this perspective has expedited advancements, it often overlooks the risk that excessively capable systems may encroach upon and ultimately supplant human cognitive functions. It also fails to adequately recognize the distinctive agency of human beings as self-aware and reflective individuals. What was initially intended as a means to offload repetitive and routine tasks has, in practice, evolved into a more complex phenomenon: driven by the allure of convenience, speed, and optimized outcomes, users increasingly outsource information acquisition, problem analysis, decision-making, and even creative tasks to AI systems. In doing so, they gradually cede their active role—transforming from originators of thought into passive validators of AI-generated outputs.
Reflecting on both the current trajectory and foundational vision of AI development, numerous concerns have been articulated on AI systems that learn and think like humans as well as the unsettling possibility that such systems may replace human roles. In response to these apprehensions, scholars and practitioners have advocated for AI systems that learn and think with humans—as auxiliary partners that augment human intelligence (Collins et al., 2024; Brynjolfsson, 2022). In other words, individuals engage with AI not merely to obtain definitive answers to their questions, but to enhance their own problem-solving abilities with its assistance (Volkman & Gabriels, 2023). Hence, rather than conceiving of AI as an omniscient, disinterested, and dispassionate ideal observer, it is preferable to activate human agency, enabling users to exercise their intellectual capacities to generate solutions tailored to their own needs. Furthermore, a genuinely human-centered AI should go beyond merely enhancing system transparency (e.g., explainable AI) or preserving meaningful user control (e.g., contestable AI) (Raees et al., 2024; Robbins, 2024). It should also aspire to cultivate human agency, encouraging users to actively engage in the operation and outcomes of AI systems and, ultimately, to foster a relationship of mutual learning, adaptation, and co-evolution between humans and AI (Ziegler & Donkers, 2024). From this perspective, the Socratic method provides both a conceptual foundation and a structural framework for addressing AI’s ethical challenges and advancing complementary human–AI collaboration.
This paper aims to bridge these identified research gaps, to elucidate autonomy risks inherent in conversational human-AI interactions, and, drawing upon existing applications of the Socratic method in AI systems (collectively termed as Socratic AI), to propose a unique Socratic chatbot, named SocrAI, as a potential solution. To this end, the paper first identifies the risks to autonomy posed by generative AI or chatbots through the lens of philosophical theories of autonomy. It then analyzes the underlying causes of user autonomy impairment and specifies the conditions a chatbot must fulfill to effectively preserve and support user autonomy. Subsequently, we discuss how the Socratic method and Socratic AI provide promising avenues for meeting these conditions. We also explore potential product formats for SocrAI and outline the further work required to actualize it. Finally, the paper concludes with a brief discussion on the significance of SocrAI in reshaping human-AI interactions and enhancing user autonomy.
Autonomy Risks in Generative Chatbots
Autonomy is generally conceptualized as the capacity to act and live in accordance with the principles of self-determination, self-governance, and self-authorship (e.g., Mackenzie, 2014; Christman, 2020). While philosophical debates continue to explore its precise nature, there is a general consensus among scholars that two foundational conditions must be met for an individual to be deemed autonomous. The first is the authenticity condition, which stipulates that one’s actions must be congruent with personal desires, beliefs, and values that genuinely reflect or constitute the “true self” (cf. Karlan, 2024). The second is competency condition, which involves a set of internal capacities that empower individuals to effectively exercise their autonomy. These capacities typically include psychological coherence, critical evaluation of information, self-awareness and reflection, rational deliberation, and the ability to formulate and revise life plans (Christman, 2009, p. 155; Roberts, 2018). Building on this conceptual framework, autonomy risks in AI systems can be broadly categorized into two main dimensions: authenticity and competency (cf. Prunkl, 2024). In the context of human–AI interaction, concerns regarding authenticity stem from the structural tension between algorithmic outputs and users’ genuine preferences, beliefs, and values. Competency, in turn, is generally understood as the agency individuals possess to support their own decision-making while interacting with AI systems. In this sense, a compromise to autonomy can be understood as a failure to effectively exercise one’s autonomous capacities, thereby preventing them from making choices and decisions that align with their authentic self.
False Mental States
In contrast to traditional AI systems, generative chatbots based on LLMs introduce a new paradigm of human–AI interaction. In this paradigm, the AI’s output does not directly determine users’ choices or decisions; rather, it influences their mental states (e.g., beliefs, values, and emotions) by providing textual information, although this influence may ultimately lead users to initiate particular actions. In this context, the central issue for the authenticity condition is whether the AI’s output provides truthful information to prevent users from adopting false mental states. This is critical because individuals’ cognitions and beliefs about themselves, society, and the world largely depend on the information they acquire (Felin & Holweg, 2024). From an epistemological standpoint, if an individual is subjected to undue external interventions and influences (e.g., coercion from social pressures or the passive acceptance and internalization of oppressive cultural norms) resulting in cognitions and beliefs to deviate from those they would endorse under semi-idealized conditions, then such beliefs can hardly be considered authentic or autonomous (Tarsney, 2025; Matheson & Lougheed, 2021). Furthermore, false beliefs (or, more broadly, false mental states) can sever the link between one’s actions and their deeper intentions, motivations, and values, preventing these actions from expressing their authentic self (Marchegiani, 2025). Viewed through this lens, the misinformation and disinformation generated by chatbots can distort users’ cognitions, leading them to make inauthentic choices and severely undermining their autonomy.
The first risk to autonomy is that misinformation can lead users to form erroneous cognitions. Since LLMs generate responses based on probabilistic frameworks such as “next word prediction,” ChatGPT is often found to produce outputs that, while grammatically and semantically plausible, may be factually inaccurate or entirely fabricated. For instance, when queried about the reasons behind Hamlet’s hypothetical actions of killing his father and marrying his mother, ChatGPT may fabricate a logically coherent yet entirely fictional explanation (cf. Sison et al., 2024). This phenomenon is commonly referred to as “hallucination.” Hallucinations generated by ChatGPT exacerbate the spread of misinformation and heighten the risk associated with AI-generated deepfakes. If users place trust in this content, they may develop distorted perceptions of reality. Should this misinformation influence real-world decision-making, the implications can be severe. In high-stakes fields such as healthcare and law, the repercussions of misinformation can be particularly grave—an inaccurate dosage recommendation could jeopardize a patient’s health, while erroneous legal advice might expose individuals to civil litigation or criminal liability (Baumgartner, 2023).
Moreover, LLMs such as ChatGPT are trained on vast datasets, the information within which may be inaccurate, misleading, or reflective of prevailing societal biases. The sheer volume of such data makes comprehensive human review impractical, which results in outputs that are liable to reproduce these errors and biases (Dai et al., 2024). For instance, empirical studies have demonstrated that ChatGPT exhibits systematic political bias, favoring entities such as the U.S. Democratic Party, Brazil’s Lula, and the U.K. Labour Party (Motoki et al., 2024). Furthermore, when certain groups, cultures, languages, or perspectives are underrepresented in the training data, AI systems tend to generalize based on dominant discourses, overlooking the uniqueness of marginalized voices. It is conceivable that when a female student consults ChatGPT regarding applications to a computer science program, the system may, influenced by historical trends, suggest that she pursue a humanities subject instead (Williams, 2024). Users who rely on chatbots for cognitive and decision-making support are thus susceptible to being misled by these erroneous and biased information. They can restrict users’ access to comprehensive and reliable information, while also reinforcing stereotypes and internalizing prejudicial values, leading to false beliefs. Critically, for underrepresented groups, ChatGPT may fail to produce positive and authentic representations that reflect their identities. Such bias can result in what has been termed “value lock-in,” systematically excluding individuals and communities that deviate from dominant societal norms and constraining the expression of personal identity (Leslie & Meng, 2024; Weidinger et al., 2022).
The second autonomy risk involves the use of disinformation to deceive and manipulate users into adopting false mental states that serve specific political or economic ends. Due to its sophisticated interactive capabilities, ChatGPT has been proposed as a vehicle for persuasive strategies designed to foster behavioral change or enhance the quality of our (moral) decision-making through “nudging” (Salvi et al., 2025; Matz et al., 2024). Yet, the use of ChatGPT also raises concerns about manipulation. Particularly in domains such as social networking, marketing, and politics that immediate returns are prevalent, rational persuasion is at risk of devolving into illicit manipulation (Klenk, 2024). Given its unprecedented capacity to generate and disseminate misleading content, ChatGPT can be exploited for political manipulation, delivering personalized misinformation to millions of users as part of targeted propaganda efforts (Tarsney, 2025). When algorithms infer users’ preferences and emotional states by analyzing their behavior and interaction histories and then target them with the most compelling products during moments of emotional vulnerability, fictitious positive reviews generated by ChatGPT may serve as the final impetus that induces a purchase (cf. Stark, 2018). In such scenarios, ChatGPT exerts a covert influence on users, causing them to adopt false mental states with the aim of advancing the specific political and economic interests of a manipulator (see Ienca, 2023). Because users remain unaware of these influences, deception and manipulation create a false sense of autonomy: users believe they are making informed decisions grounded in accurate information and personal preference, when in reality both have been subtly guided and distorted. If users were to become aware of this deception and manipulation, they would likely feel deceived and come to regret their choices and decisions.
Cognitive Deskilling
Impelled by a variety of factors, such as the need to mitigate information overload or to cope with systemic societal pressures, individuals are increasingly disposed to outsource tasks to AI systems to support their own information acquisition and decision-making (see Krook, 2025). Strategically, reliance on AI for problem-solving represents a cognitive shortcut that offers low costs and high returns. Cognitive outsourcing can not only assist individuals in completing tasks more efficiently but can also alleviate their cognitive burden and compensate for deficiencies in their knowledge and experience (Buçinca et al., 2021; Candrian & Scherer, 2022). From the standpoint of competency, however, outsourcing cognitive tasks to generative AI contributes to the degeneration of the corresponding cognitive abilities or skills. This is because the acquisition and maintenance of human skills depend on sustained practice and application. Just as muscles require exercise to maintain their strength, our cognitive skills must be continuously engaged to be consolidated and developed. Cognitive outsourcing, by contrast, fundamentally reduces—or even eliminates—opportunities to practice these essential skills. Researchers have found that individuals who outsource cognitive tasks to AI demonstrate significantly lower levels of cognitive engagement and mental effort, indicating a comprehensive takeover of their thought processes by the AI (Zhang & Xu, 2025; Lee et al., 2025). In the long term, users’ skills will gradually degrade to the point where they struggle even with tasks they were once fully competent to perform.2
Specifically, when individuals rely on search engine rankings, social media recommendations, and LLM-generated summaries to acquire information, they effectively forgo the critical processes of information screening and evaluation that are indispensable to self-governance. The process of information acquisition involves not only the retrieval of available data but also the assessment of its quality and the filtering of low-quality and irrelevant material. It is through this process that individuals are exposed to diverse viewpoints and, by means of critical analysis, form a conceptual framework for a logical and coherent understanding of the issue at hand (Rahimzadeh et al., 2023). When ChatGPT provides direct, ready-made answers, users lose the incentive to spend time selectively processing information and deriving relevant knowledge from it. With the assistance of AI, users have fewer opportunities to commit knowledge to memory, organize it logically, and internalize concepts, as well as fewer occasions to engage in critical analysis, evaluation, and rational judgment. As several studies have indicated, users who are overly reliant on AI exhibit symptoms such as memory decline, reduced concentration, and diminished analysis depth (Krook, 2025; Gerlich, 2025; Zhai et al., 2024).
Moreover, reliance on AI for information acquisition engenders a superficial comprehension of the problems at stake. Having bypassed the cognitive processes of evaluation and synthesis, users are deprived of the opportunity to construct their own cognitive frameworks through their own efforts. They can only passively accept the algorithmic output and fail to discern the connections between different viewpoints. The Q&A interaction model gradually conditions the human mind to a fragmented mode of information consumption, which severely impairs the capacity for sustained attention—a cornerstone of deep reading, complex problem analysis, and creative work. Through long-term engagement in highly homogenized question-and-answer interactions with AI, our capacity for systematic thinking and model construction is liable to diminish, leading to a “mechanical convergence” in our modes of thinking (Lee et al., 2025). Additionally, an individual’s cognitive depth is strongly dependent on the quality of information they receive. High-quality information, presented in a structured and systematic manner, can assist individuals in constructing or refining their cognitive frameworks, thereby enabling a profound understanding and grasp of a subject’s intrinsic logic and complex relations. However, the textual knowledge that LLMs generate draws from diverse yet uneven-quality sources, often lacking in accuracy and scholarly rigor. For example, in academic research, specialized knowledge is typically found in peer-reviewed journals, which are frequently inaccessible to LLMs due to paywalls. Consequently, such models rely primarily on freely available online content (Rahimzadeh et al., 2023). Under these circumstances, it is difficult for users to acquire a comprehensive and in-depth understanding of academic issues through the texts generated by AI.
Secondly, cognitive outsourcing diminishes users’ motivation and opportunities to engage in critical evaluation and reflection. Texts produced by LLMs are typically well-organized, detailed, and appear highly credible. For users lacking domain expertise, it is extremely challenging to discern which elements of the information are accurate and which are erroneous or fabricated without meticulous verification. Yet the primary purpose of cognitive outsourcing is to reduce informational and cognitive burdens; if one must then invest additional effort to verify the authenticity of AI-generated text, the initial purpose is undermined, and one might as well retrieve and organize the information independently. Consequently, users often prioritize convenience, setting aside their scrutiny of information authenticity and assuming the reliability of AI-generated outputs. Notably, when this heuristic shortcut of cognitive outsourcing consistently produces satisfactory outcomes, users become even less motivated to engage in the deliberate processes of information retrieval, comparison, and critical analysis. As a result, they are increasingly inclined to accept AI’s recommendations without independent verification due to cognitive reliance.
Moreover, in an environment where algorithms consistently provide optimal answers, opportunities for critical thinking become limited. Such thinking typically arises from exposure to diverse or conflicting perspectives and is activated when individuals encounter cognitive dissonance, contradictory information, or complex problems (Lu, 2024). However, outputs from LLMs are characteristically fluent and logically coherent, seldom presenting heterogeneity or dissenting viewpoints. Even when users harbor doubts about chatbot responses, their typical recourse is not to meticulously verify the information but to pose further questions about related details. Importantly, LLMs display a marked tendency toward sycophancy—the propensity to produce responses that align with users’ beliefs rather than objective facts. Chatbots often repeat user errors rather than correct them, even when they know the correct answer. If a user expresses uncertainty about a response, the chatbot is likely to adjust its response to accommodate the user’s belief (Sharma et al., 2025). This sycophantic bias can create an anchoring effect, whereby erroneous information is reiterated across successive interactions, gradually reinforcing users’ misconceptions and leading them to accept falsehoods as truths. When information is presented so smoothly and agreeably, users are deprived of both the opportunity and the motivation for critical thinking. Like the proverbial frog in slowly boiling water, they gradually lose their vigilance and struggle to recognize when their understanding has been subtly shaped by misinformation or to identify gaps in their own knowledge (Schneider, 2025).
From an epistemological standpoint, if autonomy is understood as the individual’s capacity to independently form, sustain, and revise their own beliefs and values, then critical reflection is a necessary condition for one to be epistemically autonomous. Critical reflection is here defined as a composite of intellectual and epistemic capacities that enable individuals to resist external interventions and control, and thus maintain self-governance. These capacities—including rational deliberation, critical thinking, and metacognition (they may overlap to varying degrees)—allow individuals to evaluate the reliability of information sources, distinguish between facts, opinions, and inferences, and ground their beliefs in sound reasons and evidence. By virtue of these capacities, individuals can become the masters of their own thoughts, rather than being subject to the sway and control of external forces.
However, as the preceding analysis has indicated, these very intellectual capacities that underpin epistemic autonomy face fundamental challenges in humans’ interaction with generative AI. First, cognitive outsourcing eliminates the crucial process through which users acquire and integrate information through their own cognitive efforts, thereby depriving them of the epistemic resources essential for critical evaluation. Second, reliance on such heuristic shortcuts diminishes users’ motivation to engage in independent evaluation and reflective thought. Third, informational environments lacking diversity further constrain opportunities for critical reflection. Furthermore, effective critical evaluation presupposes a transparent object of scrutiny, yet algorithmic opacity produces a decisional “black box” that obscures the critical causal pathways between the algorithm’s inputs and outputs. Users are thus confined to passively accepting AI-generated answers and are unable to critically scrutinize the reliability of the information sources or the validity of the reasoning process (Vaassen, 2022). If users lack sufficient algorithmic literacy and the necessary background knowledge, they are often unable to discern truth from falsehood, identify the biases therein, or recognize the covert influence of algorithms on their cognition and decision-making—let alone critically evaluate this information and the potential ethical risks that the AI may pose (see also Karlan, 2024). When both the object and conditions of critical reflection are obscured—and users lack the motivation, opportunities, and knowledge to activate reflective thinking—they become vulnerable to misinformation and manipulation by AI systems. In such cases, individuals may adopt false beliefs and make decisions misaligned with their authentic intentions, resulting in a profound erosion of personal autonomy.
Understanding and Responding to the Problem
In an ideal scenario, the most direct way to eliminate autonomy risks would be abstention—that is, refraining from the use of AI systems altogether (Karlan, 2024). When information acquisition and decision-making are undertaken solely by individuals with well-developed intellectual capacities, they can make decisions that align with their fundamental intentions and values. In this context, the risk of autonomy impairment disappears. However, humans are beings of limited cognitive capacity and finite time; it is impracticable for them to verify every piece of information independently. They inevitably rely on the discoveries and professional judgments of others, particularly within the context of the specialized and highly granular division of labor that characterizes modern knowledge. This explains why dependence on AI for cognitive and decision-making assistance has become a pervasive and seemingly irreversible trend. Even Karlan himself concedes that abstention is not the sole option, for in certain circumstances, it is neither necessary nor advisable. Given the inevitability of cognitive dependence, scholars have argued that autonomy requires “thinking for yourself” rather than “thinking by yourself” (King, 2021, p. 88). That is to say, an epistemically autonomous individual does not outright reject external influence; instead, they should selectively trust reliable sources of knowledge. In this sense, autonomy constitutes an epistemic capacity or virtue for critical self-determination, one that enables individuals to discern when to think independently and when to rely on the testimony of others (Matheson & Lougheed, 2021).
Since AI has already assumed an indispensable role in our daily lives, a more pragmatic approach is to identify the causes and mechanisms that give rise to the risks to autonomy and to formulate corresponding countermeasures. As the foregoing analysis indicates, false mental states and cognitive deskilling bring autonomy risks by respectively compromising the authenticity conditions and competency conditions necessary for its realization. False mental states arise from the misinformation and disinformation generated by chatbots, while cognitive deskilling erodes users’ capacity for critical reflection. Accordingly, two proximate causes and their mechanisms can be identified that explain how generative AI undermines user autonomy: when individuals rely on AI systems for information acquisition and decisional support, (1) cognitive deskilling prevents them from effectively exercising critical reflection so that (2) they fail to identify errors in the information and to recognize the covert influence of the AI on their own cognition and decision-making, which in turn leads them to unwittingly adopt false mental states or make decisions that contradict their authentic intentions and values, ultimately suffering a loss of autonomy. In accordance with this understanding, any effort to mitigate these risks in human–AI interaction and to support human autonomy must therefore involve (1) improving the quality of chatbot-generated content to overcome misinformation and falsehoods, and (2) cultivating users’ critical reflection capacities to encourage their systematic evaluation of AI-generated outputs.3
To achieve these two objectives, one potential strategy would be to identify the causes that lead to the generation of misinformation and disinformation by AI, as well as the factors that impede individuals from effectively exercising their capacity for critical reflection, and subsequently devise targeted countermeasures. While this approach may be regarded as remedial and problem-specific, certain subproblems resist straightforward solutions. For instance, algorithmic opacity impedes users’ capacity for critical evaluation of AI-generated content, thereby undermining their autonomy (Karlan, 2024). Yet such opacity arises from multiple, deeply rooted causes—including the intrinsic complexity of machine learning models, commercial imperatives related to intellectual property protection, and the knowledge gap that prevents ordinary users from comprehending the intricacies of algorithmic design and operation (Burrell, 2016). Overcoming these challenges is not a matter of simple solutions; at the very least, the technical remedies proposed under the rubric of Explainable and Interpretable AI (XAI) have thus far proven inadequate (cf. Saeed & Omlin, 2023; Karlan, 2024). To be sure, no single problem admits only one solution. If there were a solution capable of overcoming the limitations of such piecemeal measures—one that could systematically enhance the quality of chatbot outputs while simultaneously activating users’ critical reflection—it would offer a foundational means of mitigating the autonomy risks posed by chatbots. It is in this sense that we posit the integration of the Socratic method into AI systems as a promising and conceptually robust framework for addressing these challenges.
The Socratic method is a reflective inquiry approach that promotes critical thinking and the exploration of complex ideas and beliefs via systematic questioning (Paul & Elder, 2019, p. 4). This method is widely used across fields, including education, academic research, psychological counseling, decision-making, and structured dialogue. Common techniques within the Socratic method encompass definition, elenchus, dialectic, maieutics, generalization, induction, and causal reasoning (Chang, 2023; Overholser, 1993a, b). Due to its capacity to facilitate structured thinking guidance and dynamic knowledge verification, the Socratic method has drawn considerable attention from AI developers. Table 1 outlines its varied applications within current AI systems and their respective functions. These applications can be broadly categorized into two types: the first employes the Socratic method as an algorithmic framework to improve reasoning capabilities and output quality of LLMs (Application 1); the second aims to position generative chatbots as facilitators akin to Socrates, guiding users toward critical reflection and autonomous problem-solving through structured questioning (Application 2).
Table 1.
Overview of contemporary AI systems implementing the Socratic method
| Category | Name | Author | Function |
|---|---|---|---|
|
LLMs with the Socratic Method (Application 1: Output Optimization) |
SocraSynth | Chang (2024) | Implements a multi-agent debate framework based on the Socratic method, allowing LLM agents to engage in discussions and assess quality. |
| Automated Socratic Method | Underwood and Fenwick (2024) | Applies the Socratic method to evaluate the initial responses of LLMs and improve output quality through iterative cycles of optimization. | |
| Digital Socrates | Gu et al. (2024) | Reformulates assessment of explanation quality into Socratic explanation critique tasks, offering detailed evaluative framework and revision guidance to enhance the explanatory capabilities of LLMs. | |
|
Socratic Chatbots (Application 2: Question Guidance) |
SAMA | Lara (2021) | Simulates the role of Socrates to improve users’ moral reasoning while preserving individual autonomy through systematical questioning. |
| SocraticLM | Liu et al. (2024) | Employs a thought-provoking questioning strategy to guide students through stepwise problem-solving, supporting personalized learning experiences. | |
| Socratic Tutor | Favero et al. (2024) | Generates targeted Socratic questions tailored to specific dialogue contexts to cultivate learners’ critical thinking skills. |
One foundational form of Socratic AI is its implementation as an algorithmic framework designed to mitigate problems in LLMs such as hallucinations, algorithmic bias, and reasoning deficiencies (Application 1). Owing to its potential for dynamic knowledge verification, the Socratic method provides a philosophical framework for LLMs to transform their generation process from an opaque probabilistic sampling mechanism into an interpretable trajectory of cognitive reasoning, thereby enhancing the accuracy and quality of their outputs. In this paradigm, the Socratic method does not inject new knowledge directly into the model. Instead, it simulates Socratic dialogues that guide the LLM toward self-examination and critical reasoning, activating its intrinsic capacity for reflection and self-correction. This internal dialogue can take two forms. The first involves multi-agent debates among LLMs—one agent articulates a position, while another challenges it through questioning and rebuttal (Chang, 2024; Volkman & Gabriels, 2023). For example, Chang (2024) developed SocraSynth, a platform designed for multi-LLM reasoning based on conditional statistics. Rather than “teaching” the LLM how to reason, the platform creates an environment in which the model produces verified, high-quality knowledge through Socratic debate—two LLM agents argue from opposing positions on a given topic, continually presenting claims and counterarguments in an iterative refinement process. The second, less explicit form of Socratic dialogue involves employing the Socratic method as an evaluative model to pose structured questions to an LLM’s initial response. In this approach, the evaluation model critically examines the LLM’s outputs or explanations according to pre-established criteria. If deficiencies are detected, the evaluation model generates corresponding Socratic questions to optimize the quality of the output and, through multiple iterations, reinforces the model’s logical consistency, factual accuracy, and relevance in its content generation, thereby reducing the occurrence of hallucinations (Underwood & Fenwick, 2024; Gu et al., 2024).
Given the promising prospects of the Socratic method in collaborative dialogues, researchers have begun to explore the development of Socratic chatbots—AI systems designed to emulate the role of Socrates (Application 2). An early conceptualization of such a system is the Socratic Artificial Moral Agent (SAMA), which was proposed in response to the ethical challenges that moral enhancement technologies pose to human autonomy. SAMA is envisioned as an intelligent assistant grounded in the Socratic method, functioning as a moral tutor that assists users in producing their own moral insights and enhancing their capacity for moral reasoning and deliberation (Lara, 2021; Lara & Deckers, 2020). Rather than presenting predetermined moral doctrines or prescribing decisions, SAMA engages users through continuous dialogue and inquiry, with the goal of helping users clarify ambiguities in their conceptual understanding, critically evaluate their moral positions, develop the cognitive skills, and independently make more deliberative moral decisions (see also Volkman & Gabriels, 2023). Following the breakthroughs in LLM technology, the idea of a Socratic chatbot has gradually moved from a theoretical response to the ethics of moral enhancement to a practical application in various domains such as education, psychotherapy, and code generation (Liu et al., 2024; Gregorcic et al., 2024; Izumi et al., 2024; Chidambaram et al., 2024). Departing from the traditional “question–answer” paradigm of conventional chatbots, the Socratic chatbot introduces a new “thought-provoking” interaction model. Instead of providing ready-made answers, it acts as an epistemic midwife, engaging users in structured, inquiry-driven dialogues that guide them to generate their own insights and solutions. For example, Liu et al. (2024) proposed a thought-provoking instructional paradigm grounded in the Socratic method and designed a system called SocraticLM to deliver personalized mathematics tutoring services. Using a step-by-step guiding question decomposition strategy, SocraticLM disassembles mathematical problems into a sequence of incrementally structured sub-questions, leading students progressively toward the final answer.
Lara and his colleagues argue that integrating the Socratic method into AI systems could have the potential to mitigate the threats to personal autonomy that arise from both moral biomedical enhancement and artificial moral advisors (Giubilini & Savulescu, 2018; Savulescu & Maslen, 2015). In contrast to Lara’s position, which restricts the potential of Socratic AI in preserving personal autonomy to the domain of moral decision-making, we contend that its applicability is considerable broader. Socratic AI, instantiated as a generative chatbot, could be widely applied across diverse dialogical contexts to preserve and enhance user autonomy. Beyond serving as an algorithmic framework to improve the accuracy and reliability of LLM outputs (Application 1), Socratic AI can also operate as an interactive agent in settings that require deeper understanding and facilitate self-discovery, extending beyond education and critical thinking alone (Application 2).
Assessing the Potential of Socratic AI for Preserving and Supporting Autonomy
We now turn to examine how Socratic AI may contribute to optimizing LLM outputs and promoting users’ critical reflection (if possible), thereby mitigating the autonomy challenges posed by chatbots. With respect to the quality of LLM outputs, Socratic AI could construct a simulated environment of critical debate that compels LLMs to shift from simple information retrieval and pattern recognition to logical inference, structured argumentation, and reflective self-assessment. This mechanism—driven by external pressure (adversarial questioning) and internal reflection (self-correction)—could systematically reduce bias, eliminate hallucinations, and ultimately enhance output quality. Specifically, whether through dialogue between two LLM agents holding opposing views or through an evaluation model that poses structured questions to an LLM’s initial response, such internal dialogic mechanisms compel the model to move beyond a single-answer trajectory. For instance, SocraSynth creates an adversarial reasoning environment that encourages LLM agents to engage with divergent cultural and social perspectives, integrating them into a more comprehensive understanding of the issue (Chang, 2024). This approach could provide more nuanced and precise information, enhance the depth and breadth of argumentation, and counteract the standard Q&A model’s tendency to produce monolithic viewpoints. Moreover, critical assessment and iterative questioning could stimulate deeper reasoning processes within LLMs, prompting them to justify claims with evidence and improve the clarity of their reasoning and the intelligibility of their explanations. Over multi-round iterations, adversarial exchanges will compel the model to examine its outputs for potential inconsistencies, logical contradictions, and factual errors, as any logical leaps or factual inaccuracies are likely to be exposed by their interlocutor. This, in turn, could help LLMs progressively refine their claims, reducing the likelihood of producing unsubstantiated assertions or “hallucinations.” In a medical diagnosis case, SocraSynth was reported to have successfully identified and corrected an initial misdiagnosis, demonstrating its potential effectiveness (Chang, 2024). Empirical studies further suggest that LLMs subjected to Socratic evaluation exhibit enhanced reasoning capabilities and depth, and show significant performance improvements in coherence, factual accuracy, and relevance of their text generation (Underwood & Fenwick, 2024).
Nonetheless, a critical question remains: can Socratic AI truly help users avoid forming false beliefs stemming from algorithmic bias? Consider a hypothetical case: among humanities and arts students at American universities, there exists a well-documented correlation between major choice and gender, with females opting for literature studies and males choosing philosophy. If a predictive model were trained on this data and can relatively accurately predict the relationship between gender and major choice, should we regard this model as another stereotype or as a closer approximation of reality? Furthermore, what should be the role of Socratic AI in such a case? Should it encourage more women to pursue philosophy and more men to embrace literary studies, or the reverse?4 Indeed, the phrase “a closer approximation of reality” reflects the model’s overall predictive accuracy, while categorizing the model as “stereotyping” stems from its failure to accurately predict the choices of every individual. Given gender and major choice are not entirely congruent, the model’s inability to predict the major choices of all individuals signifies the existence of exceptional cases. For this particular case, the model’s conventional recommendations may carry a distinctly negative connotation, as they overlook individual uniqueness. Consider, for example, a female student who excels in mathematics and aspires to enroll in a computer science program. If ChatGPT were to suggest that she pursue a degree in the humanities based on statistical correlations, such a recommendation might be perceived as disappointing or even alienating, implying that her aspirations are not being recognized, but rather filtered through a stereotypical lens (see Williams, 2024).
We contend that the central issue at stake is not the statistical accuracy of the model, but whether the chatbot can present a diverse range of viewpoints. As previously argued, the activation of critical reflection depends on access to a diverse informational environment. Such an environment, ideally, would involve multiple independent and even competing sources of information that present a range of diverse and even opposing viewpoints. This exposure would afford individuals the opportunity to broadly apprehend and deeply compare different standpoints. Encountering opposing viewpoints may stimulate curiosity, encourage inquiry, and deepen understanding, while also fostering critical thinking skills (see Yamamoto, 2024). Only through being fully informed and engaging in sufficient reflection can users be expected to form sound beliefs and make choices that cohere with their values and goals.5 In this sense, Socratic AI contributes directly to viewpoint diversity by facilitating debates among two or more LLM agents that represent contrasting positions. Notably, a system such as SocraSynth is reported to generate a comprehensive report after multiple feedback loops, which includes the viewpoints of both debaters, their argumentation processes, and the results of the evaluation (Chang, 2024). Furthermore, on any given topic, Socratic AI could potentially prevent closed-mindedness by explicitly acknowledging the current limitations and uncertainties of available knowledge, as well as those aspects that remain under debate.
On another front, Socratic AI is intended to support the systematic cultivation of users’ capacity for critical reflection. In general terms, a Socratic dialogue can be conceptualized as these steps: (1) The interlocutor presents a statement (p); (2) Socrates deduces a conclusion that contradicts (p), using commonly accepted premises (q); (3) The interlocutor falls into a state of perplexity (Archie, 2010, p. 139). Through a series of strategic questions, the Socratic chatbot would likely induce cognitive tension or conflict in the user’s mind. This perplexity, arising from the recognition of internal inconsistency, may foster a sense of intellectual humility, prompting users to shift from confident assertors to cautious inquirers. In ongoing dialogue with the Socratic chatbot, users would be encouraged (or perhaps gently compelled) to examine the sources and justifications of their beliefs, provide reasons and evidence, and articulate defensible grounds for their claims. This continuous intellectual exercise could help users scrutinize their unexamined assumptions and the inconsistencies within their conceptual frameworks, thereby potentially enhancing their logical, conceptual, and empirical reasoning abilities. Over time, users may learn to approach issues from multiple perspectives and appreciate the partial validity of differing standpoints, which could lead to the development of more comprehensive and profound insights. Through the ongoing examination and revision of their own views, these dialogical techniques could become internalized. Even without the chatbot, users might then be capable of conducting a Socratic dialogue within their own minds. In other words, they would not only follow the cognitive pathways suggested by Socratic AI but also learn to think in a Socratic manner themselves—that is, to engage in self-questioning, self-justifying, and self-correcting inquiry as a habitual mode of problem-solving.6
Equipped with the capacity for critical reflection that Socratic AI could cultivate, users would be more inclined to maintain a healthy epistemic skepticism toward AI-generated content. Rather than accepting outputs at face value, they would remain aware of the covert influence of algorithms on their cognition and decision-making, as well as the potential ethical risks involved. Based on their critical evaluation, they would determine whether to trust the information provided. In certain circumstances, such users might even take proactive measures to seek out materials that algorithms “prefer” them not to see, thus breaking free from their own algorithmic filter bubbles. In engaging with information, users so trained would be less likely to accept claims simply because they align with their preexisting views or to dismiss them solely because they do not. Instead, they might adopt a stance of critical openness, prudently evaluating the plausibility and justification of competing perspectives and treating each encounter as an opportunity to test, refine, or expand their own understanding. In this way, users could escape from the algorithmic “glass cage” and gradually reclaim their agency in information consumption. From a more profound standpoint, the critical thinking skills and epistemic open-mindedness developed through interactions with Socratic chatbots are transferable beyond the realm of intellectual inquiry and can significantly enhance decision-making across various aspects of social life. Although the applications of Socratic chatbots may be more specialized than those of contemporary general-purpose generative AI—primarily focusing on domains that require deep reflection and self-discovery—this does not imply that their ability to alleviate autonomy risks is similarly limited. When users engage with other AI systems, these cultivated cognitive habits would promote a more reflective and cautious approach to AI-generated outputs.
Volkman and Gabriels (2023), however, have questioned the neutrality of the Socratic chatbot, posing that it may inadvertently reproduce the thoughts of Socrates and generate systemic biases since actual Socrates was not value-neutral but possessed his own normative commitments. Nonetheless, the training of the Socratic chatbot is not exclusively predicated on the Plato’s dialogues, but rather centers on the Socratic method as a procedural framework. In this regard, the chatbot’s emulation of Socrates is methodological rather than doctrinal—it reproduces a mode of inquiry rather than a specific set of philosophical commitments. In their interactions with the Socratic chatbot, users are to be regarded as the primary experts on their own problems. Because they possess the most direct and clear understanding and phenomenological experience of the specific context and details of their problems, and they are most aware of their own goals and needs (cf. Overholser, 1995, p. 288). Meanwhile, the Socratic chatbot would employ a methodology aligned with Schaefer and Savulescu’s (2019) concept of procedural moral enhancement. This approach would not presuppose any substantive moral principles or fixed value commitments. Instead, it would aim to assist users in organizing their thoughts and encourage them to apply their own intellectual resources to the analysis of problems, rather than thinking on their behalf. This procedural neutrality could help prevent the imposition of social biases or value-laden assumptions often embedded within algorithmic systems onto users. Consequently, the interaction between a user and the Socratic chatbot could be characterized as a process of self-initiating inquiry rather than a one-sided information transmission. In this collaborative exploration, users are not passive recipients of algorithmic content but active participants who employ their intellectual capacities in the pursuit of answers. Such a process could enhance users’ problem-solving abilities and their sense of self-efficacy.
From a more constructive standpoint, Socratic AI could potentially help mitigate the generalized responses often observed in conventional chatbots and offer a more personalized mode of interaction. Typically, when users pose questions to systems like ChatGPT, their inquiries are embedded within specific contexts rather than framed as decontextualized factual or universal questions. Effectively responding to such queries would require the chatbot to consider not only the user’s level of knowledge and intent but also to incorporate the background and context of the problem in a nuanced manner. However, existing AI models often struggle to integrate these elements effectively due to deficiencies in tracking users’ long-term conversational histories or a tendency to overlook user-specific information previously disclosed (Zhong et al., 2024). As a result, users may experience misinterpretations of intent and receive inaccurate or irrelevant responses, leading to feelings of frustration and disengagement. In contrast, Socratic AI would aim to reverse this dynamic through active, guided questioning. Such questioning could prompt the model to identify key elements, constraints, and background assumptions in the user’s problem, progressively refining the dialogue’s context. By doing so, Socratic AI might gradually apprehend the user’s intentions and prior knowledge. This dialogical fine-tuning, if realized effectively, could make subsequent exchanges more targeted and intellectually stimulating. Consequently, this personalized assistance could more effectively stimulate users’ thinking, helping them clarify ambiguous ideas, beliefs, and values, and critically reflect on their validity. For example, SocraticLM can assess a student’s cognitive state and detect misunderstandings based on their responses and pose tailored follow-up questions to help the student recognize their own errors and learning gaps, thereby supporting more effective knowledge acquisition (Liu et al., 2024). This collaborative and shared model of decision-making integrates the computational strengths of AI—particularly in processing information and integrating knowledge—with a profound respect for users’ autonomy, values, and lived experiences (Sandman & Munthe, 2009; Elwyn et al., 2012).
As discussed earlier, user autonomy is compromised during interactions with chatbots when these systems generate information that is either erroneous or fabricated for users. In the absence of critical reflection, such information can lead users to form false mental states or to make decisions that contradict their fundamental intentions. In contrast, the Socratic method employs systematic and reflexive questioning to stimulate critical reflection both within LLMs and in users themselves. These reflective processes could help models reduce internal biases and limit hallucinatory tendencies, thereby enhancing the epistemic reliability of their outputs. They may also encourage users to cultivate a healthy skepticism that promotes a more cautious and critical engagement with AI-generated content. Furthermore, by progressively refining the conversational context and prompting users to articulate their underlying assumptions and goals, Socratic AI could help individuals make choices that align more closely with their authentic intentions. In a word, these mechanisms suggest that Socratic AI—though still largely theoretical—may offer promise as a framework for mitigating autonomy risks in human–AI interaction and supporting users in exercising epistemic and practical self-governance.
The Prospect for SocrAI
Despite the considerable potential of Socratic AI in maintaining user autonomy, the realization of a fully viable product—termed SocrAI—remains a substantial challenge. It is clear that Application 1 alone is insufficient to address all autonomy risks in generative chatbots, since its strengths lie primarily in mitigating hallucinations, algorithmic bias, among others. A natural refinement involves embedding Application 1 within Application 2, wherein the chatbot adopts the role of Socrates and employs Socratic methods to enhance its outputs, thus combining the benefits of both application in preserving user autonomy. However, this integration results in a generative chatbot characterized by Socratic interaction. Such an approach may deter users, as the increased cognitive effort required could be perceived as burdensome. For users accustomed to rapid, efficient outputs from current AI systems, SocrAI may be experienced as frustratingly slow and demanding. Furthermore, users may lack the patience or willingness to engage in the often discomforting process of Socratic dialogue. Even if SocrAI can enhance autonomy, encourage self-discovery, and improve problem comprehension, the process itself can be emotionally taxing, just as Socratic dialogue frequently induces confusion, discomfort, and irritation (Boghossian, 2012). Reducing a participant’s beliefs to contradictions or absurdities may not lead to openness but rather to defensiveness and resentment (Sokoloff, 2020, pp. 61–62). Users may complain that SocrAI provokes constant doubt and undermines their psychological comfort due to being negated (cf. Airaksinen, 2022; Stoddard & O’Dell, 2016).7
An alternative design approach involves the implementation of a parallel architecture that offers users two modes within the interactive interface: a regular mode and a Socratic mode. The regular mode resembles the multi-agent framework used in SocraSynth (Chang, 2024), where multi-LLM agents articulate varied or even opposing viewpoints and refine their arguments through a process of Socratic debate. The final output synthesizes these perspectives, presenting a spectrum of ideas and supporting justifications. This structure is intended to enhance the accuracy and reliability of response while encouraging intellectual exploration by exposing users to diverse perspectives. Conversely, the Socratic mode is designed for the chatbot to emulate Socrates. It monitors the progression of dialogue, infers the user’s knowledge level, identifies their goals and beliefs, formulates questions accordingly, and adjusts its inquiry strategy in real time based on user responses. This mode is intended to cultivate critical thinking and promote users’ problem-solving abilities. From a technical perspective, a parallel architecture may be easier to implement. The development and refinement of each mode in isolation would diminish the technical complexity of model training. Moreover, the ability to toggle between modes allows users to select their preferred interaction style, thereby reducing the psychological and cognitive entry barrier. In situations where speed, efficiency, or rapid ideation are needed, users may opt for the regular mode, benefiting from optimized and diverse outputs. In contrast, when individuals seek inspiration, engage in academic inquiry, or cultivate critical thinking skills, they may transition to the Socratic mode for a more reflective and dialogical experience.
However, such an artificial separation between interaction modes may undermine the coherence and fluidity of dialogue. In the course of a continuous conversation, users may need both the provision of relevant information and the stimulation of guided inquiry. Frequent mode-switching introduces operational friction and can disrupt the user’s cognitive flow, resulting in diminished engagement. Transitioning from the regular mode to the Socratic mode necessitates that the system seamlessly integrates prior conversational context, accurately identifies the user’s intent to explore specific issues, and initiates appropriate Socratic questioning—all of which present great technical challenges. Additionally, a rigid dichotomy between modes may constrain the capabilities of SocrAI. The functionalities associated with Application 1 (optimizing outputs) and Application 2 (formulating guiding questions) are not decisively exclusive but rather intertwined. While the regular mode aims to produce high-quality answers, a truly valuable response—beyond being factually correct and contextually relevant—often includes prompts that stimulate further user reflection. Conversely, although the Socratic mode emphasizes inquiry, effective questioning presupposes a solid understanding of relevant subject matter and user background. In many cases, the chatbot must provide brief explanations or clarifications to pose meaningful questions—drawing directly on the capabilities of Application 1.
Consequently, a better design strategy may involve employing the regular mode as the underlying content-generation architecture for the chatbot, while developing an adaptive Socratic mode as an optional cognitive augmentation—similar to the optional web-browsing enhancement provided by ChatGPT. This adaptive mode would seek to seamlessly integrate Application 1 and Application 2 into a fluid dialogue system, thereby enabling SocrAI to access conversational context and user requirements in real time and respond appropriately—much like the dialogical process demonstrated in Overholser’s (1993a) therapeutic dialogue examples.
Nonetheless, the realization of this vision will need considerable further development. Currently, Application 2 is largely restricted to educational settings, particularly in the domains of mathematics, physics, and critical thinking training. The potential application of SocrAI in more complex fields such as psychotherapy, personal decision-making, and academic inquiry remains underexplored. Second, existing datasets, such as SocraTeach (Liu et al., 2024) and SocratiQ (Ang et al., 2023), are relatively simplistic and not inadequate to handle the intricate problem scenarios and varied user requirements. Third, the five categories of questions outlined by Paul and Elder (2019) do not encompass the full spectrum of Socratic questioning. Overholser (1993a), for example, catalogued various types of Socratic questions used in psychotherapy, many of which fall outside Elder and Paul’s framework. It is plausible that numerous additional, domain-specific variants of Socratic questions remain to be discovered. To enhance the contextual adaptability of SocrAI, substantial investment will be needed to develop broader and more nuanced datasets. This could involve annotating existing resources—such as philosophical dialogues, therapeutic transcripts, and teaching case studies—or collaborating with domain experts to create high-quality simulated conversations. Another viable strategy would involve the development of domain-specific fine-tuned versions of SocrAI, complete with customized question taxonomies and strategic libraries.
Furthermore, the potential discomfort that may arise due to SocrAI’s critical nature serves as a crucial reminder of the need to design the system in a way that prevents its malicious use.8 Consider a scenario wherein a user disguises an opponent’s opinion as their own, prompts SocrAI to analyze it, and then shares it responses with the opponent selectively as a seemingly neutral and subtle critique. In this scenario, SocrAI becomes a weapon for epistemic attack—not to foster understanding or truth, but to strategically destruct others’ beliefs. To guard against such misuse, it is imperative that SocrAI maintains a consistently neutral, exploratory, and nonevaluative tone. It should eschew providing “ammunition,” and instead focus on posing clarifying questions, identifying underlying assumptions, and mapping reasoning chains—avoiding content that could be weaponized for personal attacks or disparagement. Nevertheless, it is likely impossible to completely eliminate the risk of malicious use, as the ethical implications of any tool depend on how users choose to engage with it. Existing generative AI systems face similar challenges (e.g., deepfakes, disinformation, and fraud), which do not align with the developers’ original intent. Hence, the key issue is not to construct an abuse-proof system, but to strongly incentivize intended uses and cultivate users’ ethical awareness. This could involve, for instance, explicitly informing users at the beginning of their interaction with SocrAI about its core values and design purposes—honesty, openness, mutual respect, and the pursuit of truth—alongside a clear warning about the risks of malicious use.
Conclusion
The emergence of ChatGPT and similar LLMs has introduced both novel challenges and potential opportunities for human autonomy. This paper, based on the authenticity and competency conditions of autonomy, has identified two primary forms of autonomy risks posed by generative chatbots: false mental states and cognitive deskilling. These risks primarily arise from the tendency of chatbots to generate information that is false or misleading for users, along with users’ inability to effectively exercise their capacity for critical reflection to identify such errors. We have proposed that the Socratic method, understood as a technique of systematic questioning, could provide a potential theoretical framework for addressing these autonomy risks. This method may not only help reduce the hallucinations and biases in LLMs, thereby enhancing the quality of their outputs, but it could also foster users’ critical thinking, enabling them to approach AI-generated content with greater prudence. Drawing on current applications of the Socratic method in AI systems, this paper has outlined potential ways to implement SocrAI and highlighted the further work that needs to be undertaken.
Theoretically speaking, SocrAI has the potential to counteract the limitations of the Q&A paradigm of chatbots, which tends to replace human thinking, and could instead serve to fully leverage user agency. In this paradigm, the AI would no longer act as an authority that merely generates textual information or provides decision-making recommendations, but would instead function as a facilitator that guides users in their reflective processes. Correspondingly, users would become self-initiating explorers who, with SocrAI’s guidance, think through and analyze problems, seek possible solutions, and ensure that each stage of their cognition and decision-making grounded in self-determination. This shared, dialogical mode of interaction would help users make decisions more aligned with their preferences and values, and realize a form of procedural, dialogue-based cognitive enhancement. With the continued rapid development of AI technologies, it is conceivable that SocrAI could gradually become a practical reality.
Acknowledgements
The authors would like to express their heartfelt gratitude to the anonymous reviewers for their thoughtful and constructive feedback, which greatly improved the clarity and theoretical contribution of this paper. They have learned a great deal from their insightful comments and suggestions. This work could not have taken its present form without their generous intellectual engagement.
Author Contributions
Both authors contributed to the conception and development of the central argument of this paper. They jointly discussed the theoretical framework, refined the structure of the manuscript, and revised the final version. Both authors read and approved the final manuscript.
Funding
This article is funded by Moral Development Think Tank and Collaborative Innovation Center for Civic Morality and Social Ethos at Jiangsu Province, China.
Declarations
Competing Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Footnotes
For the sake of convenience, this paper sometimes uses ChatGPT as a representative example to refer to generative chatbots based on large language models (LLMs).
Recent research by Kosmyna et al. (2025) examined the neural effects of LLM-assisted writing. Compared with a “brain-only” group that completed writing tasks without external tools, the LLM group, which composed essays exclusively using GPT-4o, exhibited a 47% reduction in neural network connectivity (from 79 to 42). This suggests a negative correlation between cognitive engagement and the degree of external assistance: when the LLM performs the bulk of synthesis and ideation, the brain experiences measurable cognitive atrophy. Thus, although AI use may, in the short term, ease tasks by reducing cognitive load, over the long term, a lack of the requisite cognitive engagement and deep thought can lead to a decline in users’ cognitive and synthetic abilities (see also Bastani et al., 2024).
Admittedly, factors such as developers’ profit motives and the social biases embedded in training data also profoundly affect user autonomy. These, however, are best understood as indirect or distal factors because they affect autonomy not by direct interference with users’ cognition but through shaping the design, performance, and informational ecology of AI systems. For instance, profit-driven design may promote sycophantic or homogenized information environments and employ opaque algorithmic architectures, thereby indirectly undermining user autonomy. Extending causal analysis indefinitely to encompass all such distal influences would render the discussion unfocused and analytically diffuse.
We would express our appreciation to an anonymous reviewer for proposing this thought experiment that helps clarify the key aspects of Socratic AI.
A “sound belief” here refers to one that an individual would form under semi-idealized conditions, which need not necessarily align with factual truth. For a clarification of “semi-idealized conditions,” see Tarsney (2025).
In addition, Favero et al. (2024) developed Socratic Tutor, a system specifically designed to cultivate learners’ critical thinking skills. The authors fine-tuned a LLM using a subset of 600 context–question pairs from the SocratiQ dataset created by Ang et al. (2023), enabling the model to generate targeted Socratic questions tailored to specific dialogue contexts. Experimental results indicated that the Socratic Tutor significantly outperformed standard baseline chatbots. Furthermore, its effectiveness in promoting students’ critical thinking increased with the number of dialogue turns.
Indeed, people often exhibit extraordinary perseverance when pursuing meaningful goals. Research indicates that specific and challenging goals enhance motivation and commitment (Locke & Latham, 2002). In domains such as advanced mathematics or rigorous intellectual inquiry, difficulty is not merely an obstacle but an intrinsic component of value. Likewise, in contexts such as academic research or problem-solving, users may voluntarily engage in SocrAI’s demanding dialogue process as a path toward intellectual growth. Nevertheless, enhancing SocrAI’s adaptability, intuitiveness, and user-friendliness will substantially improve its accessibility and appeal across diverse use cases.
We are grateful to one anonymous reviewer for drawing our attention to the potential risks of the malicious use of SocrAI.
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- Airaksinen, T. (2022). Socratic irony and argumentation. Argumentation, 36(1), 85–100. 10.1007/s10503-021-09556-0 [Google Scholar]
- Ang, B. H., Gollapalli, S. D., & Ng, S. K. (2023). Socratic question generation: A novel dataset, models, and evaluation. In Proceedings of the 17th conference of the European chapter of the association for computational linguistics, (pp. 147–165). 10.18653/v1/2023.eacl-main.12
- Archie, A. M. (2010). The anatomy of a dialogue. Journal of Philosophical Research, 35, 129–146. 10.5840/jpr_2010_10 [Google Scholar]
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. SSRN. 10.2139/ssrn.4895486 [Google Scholar]
- Baumgartner, C. (2023). The opportunities and pitfalls of ChatGPT in clinical and translational medicine. Clinical and Translational Medicine, 13(3), e1206. 10.1002/ctm2.1206 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boghossian, P. (2012). Socratic pedagogy: Perplexity, humiliation, shame and a broken egg. Educational Philosophy and Theory, 44(7), 710–720. 10.1111/j.1469-5812.2011.00773.x [Google Scholar]
- Bonicalzi, S., De Caro, M., & Giovanola, B. (2023). Artificial intelligence and autonomy: On the ethical dimension of recommender systems. Topoi, 42, 819–832. 10.1007/s11245-023-09922-5 [Google Scholar]
- Botes, M. (2023). Autonomy and the social dilemma of online manipulative behavior. AI and Ethics, 3, 315–323. 10.1007/s43681-022-00157-5 [Google Scholar]
- Brynjolfsson, E. (2022). The Turing trap: The promise & peril of human-like artificial intelligence. Daedalus, 151(2), 272–287. [Google Scholar]
- Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1). 10.1145/3449287
- Burrell, J. (2016). How the machine ‘thinks’: Understanding opacity in machine learning algorithms. Big Data & Society, 3(1), 2053951715622512. 10.1177/2053951715622512 [Google Scholar]
- Candrian, C., & Scherer, A. (2022). Rise of the machines: Delegating decisions to autonomous AI. Computers in Human Behavior, 134, 107308. 10.1016/j.chb.2022.107308 [Google Scholar]
- Chang, E. Y. (2024). SocraSynth: Multi-LLM reasoning with conditional statistics (No. arXiv:2402.06634). arXiv. 10.48550/arXiv.2402.06634
- Chang, E. Y. (2023). Prompting large language models with the Socratic method. In 2023 IEEE 13th annual computing and communication workshop and conference (CCWC) (pp. 0351–0360). 10.1109/CCWC57344.2023.10099179
- Chidambaram, S., Li, L. E., Bai, M., Li, X., Lin, K., Zhou, X., & Williams, A. C. (2024). Socratic human feedback (SoHF): Expert steering strategies for LLM code generation. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Findings of the association for computational linguistics: EMNLP 2024 (pp. 15491–15502). Association for Computational Linguistics. 10.18653/v1/2024.findings-emnlp.908
- Christman, J. (2009). The politics of persons: Individual autonomy and socio-historical selves. Cambridge University Press.
- Christman, J. (2020). Autonomy in moral and political philosophy. In E. N. Zalta (Ed.), The Stanford encyclopedia of philosophy (Fall 2020 edition). https://plato.stanford.edu/archives/fall2020/entries/autonomy-moral/
- Collins, K. M., Sucholutsky, I., Bhatt, U., Chandra, K., Wong, L., Lee, M., Zhang, C. E., Zhi-Xuan, T., Ho, M., Mansinghka, V., Weller, A., Tenenbaum, J. B., & Griffiths, T. L. (2024). Building machines that learn and think with people. Nature Human Behaviour, 8(10), 1851–1863. 10.1038/s41562-024-01991-9 [Google Scholar]
- Dai, S., Xu, C., Xu, S., Pang, L., Dong, Z., & Xu, J. (2024). Bias and unfairness in information retrieval systems: New challenges in the LLM era. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining (pp. 6437–6447). 10.1145/3637528.3671458
- Del Valle, J. I., & Lara, F. (2024). AI-powered recommender systems and the preservation of personal autonomy. AI & SOCIETY, 39(5), 2479–2491. 10.1007/s00146-023-01720-2 [Google Scholar]
- Dixon, R. B. L. (2023). A principled governance for emerging AI regimes: Lessons from China, the European Union, and the United States. AI and Ethics, 3(3), 793–810. 10.1007/s43681-022-00205-0 [Google Scholar]
- Dogruel, L. (2021). What is algorithm literacy? A conceptualization and challenges regarding its empirical measurement. In M. Taddicken & C. Schumann (Eds.), Algorithms and communication (pp. 67–93). Freie Universität Berlin. 10.48541/DCR.V9.3
- Elwyn, G., Frosch, D., Thomson, R., Joseph-Williams, N., Lloyd, A., Kinnersley, P., Cording, E., Tomson, D., Dodd, C., Rollnick, S., Edwards, A., & Barry, M. (2012). Shared decision making: A model for clinical practice. Journal of General Internal Medicine, 27(10), 1361–1367. 10.1007/s11606-012-2077-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Favero, L., Pérez-Ortiz, J. A., Käser, T., & Oliver, N. (2024). Enhancing critical thinking in education by means of a Socratic chatbot (No. arXiv:2409.05511). arXiv. 10.48550/arXiv.2409.05511
- Fazelpour, S., & Danks, D. (2021). Algorithmic bias: Senses, sources, solutions. Philosophy Compass. 10.1111/phc3.12760 [Google Scholar]
- Felin, T., & Holweg, M. (2024). Theory is all you need: AI, human cognition, and causal reasoning. Strategy Science, 9(4), 346–371. 10.1287/stsc.2024.0189 [Google Scholar]
- Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1(1). 10.1162/99608f92.8cd550d1
- Fossa, F. (2024). Artificial intelligence and human autonomy: The case of driving automation. AI & SOCIETY. 10.1007/s00146-024-01955-7 [Google Scholar]
- Gabriel, I., Manzini, A., Keeling, G., Hendricks, L. A., Rieser, V., Iqbal, H., Tomašev, N., Ktena, I., Kenton, Z., Rodriguez, M., El-Sayed, S., Brown, S., Akbulut, C., Trask, A., Hughes, E., Bergman, A. S., Shelby, R., Marchal, N., Griffin, C., & Manyika, J. (2024). The ethics of advanced AI assistants (No. arXiv:2404.16244). arXiv. 10.48550/arXiv.2404.16244
- Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1). 10.3390/soc15010006
- Giubilini, A., & Savulescu, J. (2018). The artificial moral advisor. The ideal observer meets artificial intelligence. Philosophy & Technology, 31(2), 169–188. 10.1007/s13347-017-0285-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gregorcic, B., Polverini, G., & Sarlah, A. (2024). ChatGPT as a tool for honing teachers’ Socratic dialogue skills. Physics Education, 59(4), 045005. 10.1088/1361-6552/ad3d21 [Google Scholar]
- Gu, Y., Tafjord, O., & Clark, P. (2024). Digital Socrates: Evaluating LLMs through explanation critiques (No. arXiv:2311.09613). arXiv. 10.48550/arXiv.2311.09613
- Hacker, P., Engel, A., & Mauer, M. (2023). Regulating ChatGPT and other large generative AI models. In Proceedings of the 2023 ACM conference on fairness accountability and transparency (pp. 1112–1123). 10.1145/3593013.3594067
- Harrer, S. (2023). Attention is not all you need: The complicated case of ethically using large language models in healthcare and medicine. eBioMedicine, 90, 104512. 10.1016/j.ebiom.2023.104512 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Héder, M. (2023). The epistemic opacity of autonomous systems and the ethical consequences. AI & SOCIETY, 38, 1819–1827. 10.1007/s00146-020-01024-9 [Google Scholar]
- Höller, S., Dilger, T., Spiess, T., Ploder, C., & Bernsteiner, R. (2023). Awareness of unethical artificial intelligence and its mitigation measures. European Journal of Interdisciplinary Studies, 15(2), 67–89. 10.24818/ejis.2023.17 [Google Scholar]
- Hua, S., Jin, S., & Jiang, S. (2024). The limitations and ethical considerations of ChatGPT. Data Intelligence, 6(1), 201–239. 10.1162/dint_a_00243 [Google Scholar]
- Ienca, M. (2023). On artificial intelligence and manipulation. Topoi, 42(3), 833–842. 10.1007/s11245-023-09940-3 [Google Scholar]
- Izumi, K., Tanaka, H., Shidara, K., Adachi, H., Kanayama, D., Kudo, T., & Nakamura, S. (2024). Response generation for cognitive behavioral therapy with large language models: Comparative study with Socratic questioning (No. arXiv:2401.15966). arXiv. 10.48550/arXiv.2401.15966
- Jobin, A., Ienca, M., & Vayena, E. (2019). The global landscape of AI ethics guidelines. Nature Machine Intelligence, 1(9), 389–399. 10.1038/s42256-019-0088-2 [Google Scholar]
- Karlan, B. (2024). Authenticity in algorithm-aided decision-making. Synthese, 204(3), 93. 10.1007/s11229-024-04716-7 [Google Scholar]
- King, N. L. (2021). The excellent mind: Intellectual virtues for everyday life. Oxford University Press.
- Klenk, M. (2024). Ethics of generative AI and manipulation: A design-oriented research agenda. Ethics and Information Technology, 26(1), 9. 10.1007/s10676-024-09745-x [Google Scholar]
- Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X. H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task (No. arXiv:2506.08872). arXiv. 10.48550/arXiv.2506.08872
- Krook, J. (2025). When autonomy breaks: The hidden existential risk of AI. AI & SOCIETY. 10.1007/s00146-025-02397-5 [Google Scholar]
- Laitinen, A., & Sahlgren, O. (2021). AI systems and respect for human autonomy. Frontiers in Artificial Intelligence. 10.3389/frai.2021.705164. 4. [Google Scholar]
- Lara, F. (2021). Why a virtual assistant for moral enhancement when we could have a Socrates? Science and Engineering Ethics, 27, 42. 10.1007/s11948-021-00318-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Lara, F., & Deckers, J. (2020). Artificial intelligence as a Socratic assistant for moral enhancement. Neuroethics, 13(3), 275–287. 10.1007/s12152-019-09401-y [Google Scholar]
- Lee, H. P. (Hank), Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–22. 10.1145/3706598.3713778
- Leslie, D., & Meng, X. L. (2024). Future shock: Grappling with the generative AI revolution. Harvard Data Science Review, Special Issue 5. 10.1162/99608f92.fad6d25c
- Liu, J., Huang, Z., Xiao, T., Sha, J., Wu, J., Liu, Q., Wang, S., & Chen, E. (2024, November 6). SocraticLM: Exploring Socratic personalized teaching with large language models. In The thirty-eighth annual conference on neural information processing systems. https://openreview.net/forum?id=qkoZgJhxsA
- Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation: A 35-year odyssey. American Psychologist, 57(9), Article 9. 10.1037/0003-066X.57.9.705
- Lu, W. (2024). Inevitable challenges of autonomy: Ethical concerns in personalized algorithmic decision-making. Humanities and Social Sciences Communications, 11. 10.1057/s41599-024-03864-y
- Mackenzie, C. (2014). Three dimensions of autonomy: A relational analysis. In A. Veltman, & M. Piper (Eds.), Autonomy, oppression, and gender (pp. 15–41). Oxford University Press. 10.1093/acprof:oso/9780199969104.003.0002
- Marchegiani, B. (2025). Anthropomorphism, false beliefs, and conversational: How chatbots undermine users’ autonomy. Journal of Applied Philosophy, japp.70008. 10.1111/japp.70008
- Matheson, J., & Lougheed, K. (2021). Introduction: Puzzles concerning epistemic autonomy. In Epistemic autonomy. Routledge.
- Matz, S. C., Teeny, J. D., Vaid, S. S., Peters, H., Harari, G. M., & Cerf, M. (2024). The potential of generative AI for personalized persuasion at scale. Scientific Reports, 14(1), 4692. 10.1038/s41598-024-53755-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mökander, J., Schuett, J., Kirk, H. R., & Floridi, L. (2024). Auditing large language models: A three-layered approach. AI and Ethics, 4(4), 1085–1115. 10.1007/s43681-023-00289-2 [Google Scholar]
- Motoki, F., Neto, P., V., & Rodrigues, V. (2024). More human than human: Measuring ChatGPT political bias. Public Choice, 198(1), 3–23. 10.1007/s11127-023-01097-2 [Google Scholar]
- Nyholm, S. (2024). Artificial intelligence and human enhancement: Can AI technologies make us more (artificially) intelligent? Cambridge Quarterly of Healthcare Ethics, 33(1), 76–88. 10.1017/S0963180123000464 [DOI] [PubMed] [Google Scholar]
- Oniani, D., Hilsman, J., Peng, Y., Poropatich, R. K., Pamplin, J. C., Legault, G. L., & Wang, Y. (2023). Adopting and expanding ethical principles for generative artificial intelligence from military to healthcare. Npj Digital Medicine, 6(1), 225. 10.1038/s41746-023-00965-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Overholser, J. C. (1993a). Elements of the Socratic method: I. Systematic questioning. Psychotherapy: Theory Research Practice Training, 30(1), 67–74. 10.1037/0033-3204.30.1.67 [Google Scholar]
- Overholser, J. C. (1993b). Elements of the Socratic method: II. Inductive reasoning. Psychotherapy: Theory Research Practice Training, 30(1), 75–85. 10.1037/0033-3204.30.1.75 [Google Scholar]
- Overholser, J. C. (1995). Elements of the Socratic method: IV. Disavowal of knowledge. Psychotherapy: Theory Research Practice Training, 32(2), 283–292. 10.1037/0033-3204.32.2.283 [Google Scholar]
- Paul, R., & Elder, L. (2019). The thinker’s guide to Socratic questioning. Rowman & Littlefield.
- Prunkl, C. (2024). Human autonomy at risk? An analysis of the challenges from AI. Minds and Machines, 34(3), 26. 10.1007/s11023-024-09665-1 [Google Scholar]
- Raees, M., Meijerink, I., Lykourentzou, I., Khan, V. J., & Papangelis, K. (2024). From explainable to interactive AI: A literature review on current trends in human-AI interaction. International Journal of Human-Computer Studies, 189, 103301. 10.1016/j.ijhcs.2024.103301 [Google Scholar]
- Rahimzadeh, V., Kostick-Quenet, K., Barby, B., J., Amy, L., & McGuire (2023). Ethics education for healthcare professionals in the era of ChatGPT and other large language models: Do we still need it? The American Journal of Bioethics, 23(10), 17–27. 10.1080/15265161.2023.2233358 [Google Scholar]
- Robbins, S. (2024). The many meanings of meaningful human control. AI and Ethics, 4(4), 1377–1388. 10.1007/s43681-023-00320-6 [Google Scholar]
- Roberts, J. T. F. (2018). Autonomy, competence and non-interference. Hec Forum, 30(3), 235–252. 10.1007/s10730-017-9344-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Saeed, W., & Omlin, C. (2023). Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities. Knowledge-Based Systems, 263, 110273. 10.1016/j.knosys.2023.110273 [Google Scholar]
- Salvi, F., Horta Ribeiro, M., Gallotti, R., & West, R. (2025). On the conversational persuasiveness of GPT-4. Nature Human Behaviour, 9(8), 1645–1653. 10.1038/s41562-025-02194-6 [Google Scholar]
- Sandman, L., & Munthe, C. (2009). Shared decision-making and patient autonomy. Theoretical Medicine and Bioethics, 30(4), 289–310. 10.1007/s11017-009-9114-4 [DOI] [PubMed] [Google Scholar]
- Savulescu, J., & Maslen, H. (2015). Moral enhancement and artificial intelligence: Moral AI? In J. Romportl, E. Zackova, & J. Kelemen (Eds.), Beyond artificial intelligence: The disappearing human-machine divide (pp. 79–95). Springer. 10.1007/978-3-319-09668-1_6
- Schaap, G., Bosse, T., & Hendriks Vettehen, P. (2024). The ABC of algorithmic aversion: Not agent, but benefits and control determine the acceptance of automated decision-making. AI & SOCIETY, 39, 1947–1960. 10.1007/s00146-023-01649-6 [Google Scholar]
- Schaefer, G. O., & Savulescu, J. (2019). Procedural moral enhancement. Neuroethics, 12(1), 73–84. 10.1007/s12152-016-9258-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schneider, S. (2025). Chatbot epistemology. Social Epistemology, 39(5), 570–589. 10.1080/02691728.2025.2500030 [Google Scholar]
- Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2025). Towards understanding sycophancy in language models (No. arXiv:2310.13548). arXiv. 10.48550/arXiv.2310.13548
- Sison, A. J. G., Daza, M. T., Gozalo-Brizuela, R., & Garrido-Merchán, E. C. (2024). ChatGPT: More than a weapon of mass deception ethical challenges and responses from the human-centered artificial intelligence (HCAI) perspective. International Journal of Human–Computer Interaction, 40(17), 4853–4872. 10.1080/10447318.2023.2225931 [Google Scholar]
- Sokoloff, W. W. (2020). Against the Socratic method. In W. W. Sokoloff (Ed.), Political science pedagogy: A critical, radical and utopian perspective (pp. 51–68). Springer. 10.1007/978-3-030-23831-5_3
- Stahl, B. C., Antoniou, J., Ryan, M., Macnish, K., & Jiya, T. (2022). Organisational responses to the ethical issues of artificial intelligence. AI & SOCIETY, 37, 23–37. 10.1007/s00146-021-01148-6 [Google Scholar]
- Stark, L. (2018). Algorithmic psychometrics and the scalable subject. Social Studies of Science, 48(2), 204–231. 10.1177/0306312718772094 [DOI] [PubMed] [Google Scholar]
- Steinerová, J. (2023). Ethical issues of human information behaviour and human information interactions. Open Information Science, 7(1), 20220155. 10.1515/opis-2022-0155 [Google Scholar]
- Steyvers, M., & Kumar, A. (2024). Three challenges for AI-assisted decision-making. Perspectives on Psychological Science, 19(5), 722–734. 10.1177/17456916231181102 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stoddard, H. A., & O’Dell, D. V. (2016). Would Socrates have actually used the Socratic method for clinical teaching? Journal of General Internal Medicine, 31(9), 1092–1096. 10.1007/s11606-016-3722-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tarsney, C. (2025). Deception and manipulation in generative AI. Philosophical Studies. 10.1007/s11098-024-02259-8 [Google Scholar]
- Tiribelli, S., & Calvaresi, D. (2024). Rethinking health recommender systems for active aging: An autonomy-based ethical analysis. Science and Engineering Ethics, 30(3), 22. 10.1007/s11948-024-00479-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Underwood, H., & Fenwick, Z. (2024). Implementing an automated Socratic method to reduce hallucinations in large language models. OSF. 10.31219/osf.io/grh4b [Google Scholar]
- Vaassen, B. (2022). AI, opacity, and personal autonomy. Philosophy & Technology, 35(4), 88. 10.1007/s13347-022-00577-5 [Google Scholar]
- Volkman, R., & Gabriels, K. (2023). AI moral enhancement: Upgrading the socio-technical system of moral engagement. Science and Engineering Ethics, 29, 11. 10.1007/s11948-023-00428-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P. S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., et al. (2022). Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM conference on fairness accountability and transparency (pp. 214–229). 10.1145/3531146.3533088
- Williams, R. T. (2024). The ethical implications of using generative chatbots in higher education. Frontiers in Education, 8. 10.3389/feduc.2023.1331607
- Yamamoto, Y. (2024). Suggestive answers strategy in human-chatbot interaction: A route to engaged critical decision making. Frontiers in Psychology, 15, 1382234. 10.3389/fpsyg.2024.1382234 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhai, C., Wibowo, S., & Li, L. D. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: A systematic review. Smart Learning Environments, 11(1), 28. 10.1186/s40561-024-00316-7 [Google Scholar]
- Zhang, L., & Xu, J. (2025). The paradox of self-efficacy and technological dependence: Unraveling generative impact on university students’ task completion. The Internet and Higher Education, 65, 100978. 10.1016/j.iheduc.2024.100978 [Google Scholar]
- Zhong, W., Guo, L., Gao, Q., Ye, H., & Wang, Y. (2024). MemoryBank: Enhancing large language models with long-term memory. Proceedings of the AAAI Conference on Artificial Intelligence, 38(17), 19724–19731. 10.1609/aaai.v38i17.29946 [Google Scholar]
- Zhou, J., Müller, H., Holzinger, A., & Chen, F. (2024). Ethical ChatGPT: Concerns, challenges, and commandments. Electronics, 13(17), 3417. 10.3390/electronics13173417 [Google Scholar]
- Ziegler, J., & Donkers, T. (2024). From explanations to human-AI co-evolution: Charting trajectories towards future user-centric AI. I-Com, 23(2), 263–272. 10.1515/icom-2024-0020 [Google Scholar]
