Abstract
Interactive intelligent agents are being integrated across society. Despite achieving human-like capabilities, humans’ responses to these agents remain poorly understood, with research fragmented across disciplines. We conducted a systematic synthesis comparing a range of psychological and behavioural responses in matched human-agent vs. human-human dyadic interactions. A total of 162 eligible studies (146 contributed to the meta-analysis; 468 effect sizes) were included in the systematic review and meta-analysis, which integrated frequentist and Bayesian approaches. Our results indicate that individuals exhibited less prosocial behaviour and moral engagement when interacting with agents vs. humans. They attributed less agency and responsibility to agents, perceiving them as less competent, likeable, and socially present. In contrast, individuals’ social alignment (i.e., alignment or adaptation of internal states and behaviours with partners), trust in partners, personal agency, task performance, and interaction experiences were generally comparable when interacting with agents vs. humans. We observed high effect-size heterogeneity for many subjective responses (i.e., social perceptions of partners, subjective trust, and interaction experiences), suggesting context-dependency of partner effects. By examining the characteristics of studies, participants, partners, interaction scenarios, and response measures, we also identified several moderators shaping partner effects. Overall, functional behaviours and interaction experiences with agents can resemble those with humans, whereas fundamental social attributions and prosocial/moral concerns lag in human-agent interactions. Agents are thus afforded instrumental value on par with humans but lack comparable intrinsic value, providing implications for the development of interactive intelligent agents.
Subject terms: Human behaviour, Computational science
This systematic review and meta-analysis found that functional behaviours and interaction experiences with intelligent agents can resemble those with humans, whereas social attributions and prosocial/moral concerns lag in human-agent interactions.
Introduction
Interactive intelligent agents, such as chatbots, virtual humans, and robots, are increasingly embedded across professional and personal domains, including healthcare, education, business, and leisure1. Recent breakthroughs in artificial intelligence (AI) have advanced these systems towards, and in some cases beyond, human-level performance on specific tasks2–5. Leading technology labs like Google’s DeepMind suggest no inherent technical limit to these systems achieving human-level capabilities across various domains6. Evolving from narrow applications to broader competence, these systems move beyond reactive “tools” or “assistants” to proactive “agents”: autonomous computational systems capable of goal-directed behaviour, interaction with the environment, and task execution with minimal human involvement7–10. Human interaction with these agents fundamentally differs from traditional human-computer interaction11. Agents are assuming roles primarily reserved for humans, serving as companions, collaborators, advisors, and even opponents12–15. Scenarios such as cooperation, competition, and strategic encounter, originally confined to human-human interaction, now emerge in human-agent interaction16–19. These developments are blurring long-standing boundaries between humans and machines.
As people engage with intelligent agents in socially complex settings requiring psychological attunement20–22, a key question emerges: can agents’ human-like performance evoke the same psychological and/or behavioural responses—the way individuals feel, think and act within an interaction23—as those with human partners? This knowledge gap calls for rigorous investigation. For example, individuals may or may not maintain comparable task performance when collaborating with agents as they do with humans. Their sense of personal agency may vary during agent interaction. Core social responses, such as alignment, trust, and morality, may be inherently tied to human presence or may be elicited by agents displaying human-like behaviour. Insights into the transferability of human psychology and behaviour from human-human interaction to human-agent interaction are critical to inform responsible agent design and regulation24–27. It is thus important to compare a range of psychological and behavioural responses in human-agent vs. human-human interactions, where the task performance of agent and human partners is matched, and to further explore the conditions under which similarities and differences emerge.
Empirical evidence on human responses in human-agent vs. human-human interactions is fractured across disciplinary silos and methodological traditions. The rise of interactive intelligent agents has attracted scholars beyond computer science and robotics, including psychology, communication, business, and marketing, spurring a proliferation of related studies28. Existing research exhibits inconsistent constructs, divergent methodologies, and isolated investigations. Specifically, similar underlying constructs have been operationalised using different indicators and measures. Taking trust towards partners as an example, it has been assessed using multiple indicators, including perceived trustworthiness, trust intention, self-reported trust, and behavioural measures29–32. Even when nominally assessing the same indicator, studies use various instruments that capture different aspects and levels of abstraction33. Research paradigms also vary considerably. Many studies examine participant responses through direct interaction with agent or human partners or through vignette-based designs in which participants imagine interacting with partners and report hypothetical responses34–36. Other studies rely on passive evaluations of partners or their generated outputs within an “interaction” framing37,38. Varying experimental control, such as mismatched partner behaviour or task performance, can further confound the interpretation of response comparability39–41. Moreover, many investigations remain isolated as one-off demonstrations, with limited follow-up, replication, or systematic extension to establish cumulative evidence42,43. These practices have contributed to largely mixed empirical findings44,45. For instance, some studies report no significant difference in subjective trust between agent and human partners46, whereas others reveal significant differences47 or qualitatively distinct trust responses48. It remains unclear whether such inconsistencies stem from research designs, agent types, interaction tasks, or participant demographics.
The mixed empirical evidence is mirrored by a diversity of theoretical perspectives. For example, the Computers Are Social Actors (CASA) paradigm and its predecessor, the Media Equation, propose that individuals respond to intelligent agents with minimal human-like cues as they would to humans, suggesting that social responses in human-agent interaction mirror those in human-human interaction49,50. This proposition of equivalence, however, is challenged by other frameworks. The Threshold Model of Social Influence draws a key distinction: when behaviour appears equally realistic, individuals show stronger deliberate social responses when they believe they are with humans rather than agents in virtual settings—though automatic, low-level reactions remain unchanged across both51. Theories of anthropomorphism shift focus to the attribution of human-like characteristics to agents’ real or imagined behaviour52,53. It has been suggested that anthropomorphising behaviour (i.e., the observable ways in which individuals respond to agents as they would to humans) should be studied in human-agent interaction without presuming response equivalence to human-human interaction54. Theories focusing on social perceptions of agents provide additional insights. For example, the Modality-Agency-Interactivity-Navigability (MAIN) model proposes that perceived agency (human vs. algorithm) activates different cognitive heuristics influencing perceived credibility55. The Social Presence Theory examines the degree to which agents are perceived as real social actors56, with a meta-analysis showing how social cues in agents increase perceived social presence57. The Uncanny Valley Theory introduces a critical caveat regarding human-like design, warning that near-human realism with subtle imperfections evokes eeriness and discomfort58. This abundance of existing theoretical frameworks underscores the need for a cross-disciplinary, systematic understanding of human responses in human-agent vs. human-human interactions, and for identifying factors influencing their similarities or differences.
Five previous reviews have addressed related topics. One narrative review argued that humans engage with agents in ways resembling interaction with other humans, including relationship building, suggesting that similarities outweigh differences and that theories of human-human interaction apply to human-agent interaction25. By contrast, another narrative review cautioned against equating these two interactions, arguing that current agents lack key social affordances underpinning theories of human-human interaction and urging the development of models specific to human-agent interaction27. An integrative review comparing trust across both interactions proposed similar developmental processes but differences in expression and calibration, and suggested narrowing gaps between perceptions of humans and agents to improve user trust59. Nevertheless, these reviews have two primary limitations. They prioritised argument-driven theoretical integration over evidence-driven systematic synthesis. Also, they lacked focus on literature directly comparing responses between these two interactions, which, even when revealing response similarities, only metaphorically support the Media Equation60.
One meta-analysis systematically compared responses in human-agent and human-human interactions by specific partner types (virtual agents vs. avatars) and found that avatars exerted stronger social influence than agents45. While insightful, this review collapsed responses (e.g., subjective perceptions, affect, task performance, physiological measures) into uniform so-called measures of social influence and pooled them in one meta-analysis. Although most psychological and behavioural responses in human-agent and human-human interactions fall under the umbrella of social responses61, treating them as a monolithic entity obscures conceptual distinctions. Another meta-analysis similarly synthesised diverse responses under the construct of persuasion and found that AI agents did not significantly differ from humans in overall persuasion outcomes44. Its operationalisation of persuasion encompassed a broader range of responses than persuasion in the conventional sense. For example, some included studies assessed evaluative or experiential responses (e.g., trustworthiness, credibility, pleasantness, affinity, and communication competence)62–65. In addition, many responses were measured in paradigms where participants evaluated AI- or human-generated outputs without engaging in interaction with a communicative partner. Accordingly, existing reviews have yet to offer a holistic picture of how various types of human responses compare between human-agent and human-human interactions.
The current review
To address prior limitations, we conducted a systematic review and meta-analysis of individuals’ psychological and behavioural responses in dyadic interactions with agent vs. human partners. The partners were functionally equivalent (i.e., performance-matched) to ensure comparability. We quantified the effects of partner type (agent vs. human; hereafter, partner effects) on specific responses and explored moderators of these partner effects. Specifically, we aimed to answer four research questions:
RQ1: Which psychological and behavioural responses have been investigated in studies comparing human-agent and human-human interactions?
RQ2: Which specific response types differ between human-agent and human-human interactions?
RQ3: Which specific response types are similar between human-agent and human-human interactions?
RQ4: To what extent do study, participant, partner, interaction, and response characteristics moderate partner effects on different response types?
As discussed above, multiple theoretical perspectives seek to explain human responses to intelligent agents. These perspectives offer overlapping yet partially conflicting predictions regarding the extent to which humans engage with agents in ways resembling interaction with other humans. Consequently, the current review employed a theory-integrative approach without relying on a single deductive theoretical lens. Rather than presuming uniform equivalence or systematic divergence, we treated response comparability as an empirical question to be investigated across psychological and behavioural constructs. In addition to quantifying the extent of similarity and difference across distinct response types, we also aimed to identify empirical boundary conditions under which theoretical predictions of equivalence or divergence are supported.
Aligning with a human-centred perspective, we treated “agent” as a functional metaphor rather than a strict technical term7. A computational system was considered an intelligent agent if it assumed a human-equivalent role while matching the task performance of its human counterpart, irrespective of technical implementation. Agent architectures may include traditional machine learning, rule-based algorithms, Wizard-of-Oz setups, or more recent generative AI. In addition, the term “interaction” is central to human-computer and human-agent interaction research, yet remains overloaded and ambiguous66. It has been understood in various ways: as an experiential stream of subjective expectations, experiences, and memories67; as a system’s disposition for interaction from a design perspective68,69; or as a process involving mutual exchanges70. Notably, we define interaction as the engagement between participants and partners for some purpose within certain contexts, involving reciprocal information exchange or unidirectional transmission with at least one party taking an active role. This requires participatory engagement, where participants act as interactants rather than observers in real-time or hypothetical (i.e., vignette-based imaginative) scenarios. Studies in which participants only evaluated AI or human partners, or their generated content, were not considered to involve interaction. Given critiques of understanding interaction as mutual exchanges in human-computer interaction66, we also included unidirectional interactions—for example, participants speak while partners listen or vice versa—provided partners are not presented as passive prerecorded stimuli but possessing interactability (i.e., the disposition and readiness to engage)71.
Furthermore, mixed empirical evidence on human responses in human-agent vs. human-human interactions implies the existence of moderating effects. We chose potential moderators based on prior related meta-analyses examining heterogeneity in human responses. Meta-analytic work comparing virtual agents and avatars in social influence examined moderators such as agency operationalisation (human- or algorithm-controlled), task nature (cooperative, competitive, or neutral), and response measure (subjective or objective)45. Another work comparing AI and human communicators in persuasion examined moderators related to participant sociodemographics, study setting (lab, online, or field), communication direction (unidirectional or bidirectional), and response domain (behaviour, perception, attitude, or intention)44. Four meta-analyses on the effects of agent anthropomorphism and social cues on human responses examined characteristics of studies (e.g., study setting, publication year), samples (e.g., age, gender, location), agents (e.g., appearance, voice, embodiment, physical or virtual form), and contexts (e.g., task type, structured or unstructured interaction)57,72–74. Drawing on this body of work and the Media Are Social Actors (MASA) paradigm, which posits that agent features, individual differences, and contextual factors shape responses to agents75, we explored moderators spanning study, participant, partner, interaction, and response characteristics. The conceptual framework of the current review is visualised in Fig. 1.
Fig. 1. Review conceptual framework.
Partner type (agent vs. human) was examined as the independent variable. Psychological and behavioural responses extracted from individual studies were classified a posteriori into distinct response types and meta-analysed separately. Partner effects were quantified as standardised mean differences (Hedges’ g) between human-agent and human-human interactions for different response types. Five categories of potential moderators were examined to explore sources of heterogeneity in partner effects.
Methods
Literature search and eligibility criteria
This systematic review and meta-analysis followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 statement. It was not preregistered and had no formal protocol. The PRISMA checklist is available in Supplementary Table 1. We conducted systematic searches in February 2024 across four literature databases: Scopus (which covers IEEE Xplore76), Web of Science, ACM Digital Library, and PsycInfo. We searched titles, abstracts, and keywords using search strings comprising two components. The first component identified various types of interactive intelligent agents, while the second combined general and specific terms to capture a wide range of topics in human-human interaction research, including common interpersonal and social phenomena that have been studied in human-agent interaction. The searches focused on empirical research articles published from 2000 and employing quantitative or mixed methods (extracting only quantitative data) and were restricted to peer-reviewed journal or conference papers written in English. The full search strings can be found in Supplementary Table 2. Grey literature was not searched due to challenges in locating studies and the lack of a reliable method for assessing quality.
We defined the eligibility criteria via the PECO (Population, Exposure, Comparison, and Outcome) framework. Eligible studies were on healthy adult participants (P) that investigated both dyadic human-agent (E) and human-human (C) interactions and made direct, parallel comparisons of human responses (O) in these two conditions. Interactive intelligent agents could be either physical or virtual, embodied or disembodied. Participants actively engaged in research tasks that involved (perceived) real-time or hypothetical interactions with partners, whom participants believed to be either an agent or human. Additionally, participants received comparable treatment across human-agent and human-human interaction conditions; that is, the interaction task, dynamics, and partner role and performance were consistent across both conditions, except for the partner’s voice and appearance. When voice and appearance were matched between agent and human partners, the only difference lay in the perceived identity of the interaction partner, isolating the identity effects on human responses. However, as voice and appearance were often used to enhance the believability of the manipulated partner identity, variations in these two aspects were permitted. Lastly, human responses were examined in terms of participants’ individual-level psychological and behavioural responses during or after the interactions.
An example study77 is presented to illustrate the type of studies included in this review. In this study, participants delivered an impromptu speech and received feedback from either a robot or a human evaluator. To ensure functional equivalence, evaluator behaviour was matched across conditions: both robot and human evaluators followed an identical script consisting of praise and constructive feedback. Moreover, the human evaluator practiced the feedback delivery to match the robot’s pacing and animacy, and the physical height of both evaluators was matched. Following the interaction, participants reported their perceptions of the evaluator and the feedback received. Such a parallel experimental design enables the examination of partner effects by manipulating partner type while ensuring comparable participant treatment.
The detailed inclusion and exclusion criteria are provided in Supplementary Table 3. The study selection involved two stages: title-abstract screening and full-text screening. At each stage, the first author screened all the candidate articles, and two co-authors each independently screened a random 15% subset78. Any discrepancies were resolved through discussion.
Data extraction and coding
The data extraction and coding protocol (including study details and essential statistics for effect size calculation; Supplementary Table 4) was established a priori, except that response classification and the specific levels of two moderators (i.e., interaction task and response dimension) were determined post hoc. The first author completed data extraction and coding for all eligible studies, and two co-authors each checked a random 15% subset. Any discrepancies were resolved through discussion.
Human responses examined across studies were diverse. Prior meta-analytic work has synthesised these responses under umbrella constructs (i.e., social influence45 and persuasion44), or grouped them into broad categories such as attitudes, perceptions, affect, and behaviour73,74. While informative, such categorisations may mask meaningful differences between conceptually distinct response types. Therefore, we performed an a posteriori classification and meta-analysed each response type separately. We initially applied a conceptual-to-empirical approach79, but the extracted responses did not align with the responses collected in Krpan’s taxonomy23. We thus shifted to an empirical-to-conceptual approach79, inductively classifying responses into distinct types and then iteratively grouping conceptually aligned response types into overarching themes. The response classification was completed before any meta-analyses commenced. The classification was primarily conducted by the first author, who has an interdisciplinary background in psychology, human-computer interaction, and AI-based system design. The process was supervised by two senior co-authors with expertise in psychology and human factors engineering, each of whom reviewed a random 15% subset to ensure classification reliability. In addition, we identified five categories of potential moderators to explore sources of heterogeneity in effect sizes: study characteristics, participant characteristics, partner characteristics, interaction characteristics, and response characteristics. The specific moderator variables and their coding scheme are detailed in Table 1.
Table 1.
Moderator variables and coding scheme
| Category | Moderator | Definition | Coding levels |
|---|---|---|---|
| Study characteristics | Study setting | Environment where the study was conducted | Laboratory; online; field |
| Study design | Assignment of participants to human-agent and human-human interaction conditions | Between-subjects; within-subjects | |
| Publication year | Year the study was published | Continuous | |
| Publication type | Type of outlet where the study was published | Journal; conference | |
| Participant characteristics | Sample continent | Continent where the study sample was located | Africa; Asia; Europe; North America; Oceania; South America |
| Sample WEIRD | Whether the study sample was in a WEIRD or non-WEIRD country | Non-WEIRD; WEIRD | |
| Mean age | Mean age of participants | Continuous | |
| Percentage female | Percentage of female participants | Continuous | |
| Partner characteristics | Human partner type | Type of human partner with whom participants interacted | Research team member (experimenter, research assistant, or hired actor/specialist following specific interaction scripts); participant (untrained individual); pseudo-human (ostensibly human partner controlled by algorithms); vignette-described partner (human partner and their behaviour described through vignettes) |
| Agent operationalisation | How the agent partner’s behaviour was operationalised during the interaction | Autonomous (behaving independently via algorithms); Wizard-of-Oz (partially or fully controlled by hidden human operators); vignette-described | |
| Agent form | Whether the agent partner existed in virtual or physical form | Virtual (existing within digital interfaces); physical (tangible in the physical world) | |
| Agent embodiment | Whether the agent partner had a visible embodied representation | Disembodied (no visible representation, e.g., bots); embodied (visibly represented, e.g., robots and virtual humans) | |
| Robot appearance | Degree of human-likeness in the robot’s appearance | Non-humanoid (ABOT score 0–10); semi-humanoid (10–40); humanoid (40–70); android (70–100) | |
| Appearance difference | Whether the agent and human partners differed in appearance | Matched; differed | |
| Voice difference | Whether the agent and human partners differed in voice | Matched; differed | |
| Interaction characteristics | Interaction realism | Whether the interaction was real-time or hypothetical | Real-time (actual or perceived real-time engagements with partners); hypothetical (vignette-based imaginative interactions) |
| Interaction flow | Whether the interaction involved bidirectional or unidirectional information exchange | Bidirectional (mutual exchanges with partners); unidirectional (one-sided input) | |
| Interaction medium | Channel through which the interaction occurred | Face-to-face (physically co-present interaction with partners); computer-mediated (interaction with partners via a computer interface without physical co-presence); virtual reality-mediated (interaction within virtual reality environments) | |
| Interaction structure | Extent to which the interaction followed a predefined or scripted format | Non-structured; semi-structured; structured | |
| Interaction nature | Goal alignment and relational dynamics shaping exchanges between participants and partners | Neutral; cooperative; oppositional; mixed | |
| Interactant power symmetry | Relative power or status balance between participants and partners | Symmetrical (equal power or status, e.g., fellow game players); asymmetrical (power imbalance, e.g., customer–service employee dyad) | |
| Interaction task | Primary task or activity characterising the interaction | Service encounter (e.g., customer service); game interaction (e.g., economic games), instructional interaction (e.g., tutoring); communication-focused interaction (e.g., dialogues and question-answer exchanges without explicit service provision, game mechanics, or instructional objectives); motor task (e.g., physical coordination); other (unclassified) | |
| Response characteristics | Response domain | Domain of participant responses measured in the interaction | Psychological (subjective perceptions and experiences); behavioural (observable behaviours and performance metrics) |
| Measurement timing | Timing of response measurement relative to the interaction | During (e.g., in-task behaviour); after the interaction (e.g., post-interaction scale) | |
| Response dimension | Specific dimensions (subconstructs) of the response type | Varied by response type |
Code as NA (not available) when relevant information is not reported or falls outside the predefined variable levels. Publication year was used as a proxy continuous moderator to explore potential temporal trends in partner effects, as few studies reported data collection time. Sample continent coding included Africa and South America, but no studies from these continents were identified. For sample WEIRD, WEIRD represents Western, Educated, Industrialised, Rich, and Democratic130. Robot appearance was coded for studies involving robots as agent partners and classified into four human-likeness levels based on the Anthropomorphic roBOT (ABOT) Database/Predictor152. Unlike other moderators with predefined levels, interaction task was coded inductively based on emergent themes observed across studies. Similarly, response dimension was coded where applicable; some response types were further classified into different dimensions (i.e., subconstructs), and this variation was tested as a post hoc moderator.
Research quality assessment
We assessed the research quality of each eligible study, as its quality can bias effect sizes and higher-quality studies generally yield findings that more closely converge on the truth80. The first author completed quality assessment for all studies, and two co-authors each checked a random 15% subset. Any discrepancies were resolved through discussion.
Existing quality assessment tools, however, posed challenges. Widely used ones, such as the Cochrane Risk of Bias 2 (RoB-2)81 and Risk of Bias in Systematic Reviews (ROBIS)82, were developed for health-related intervention research and tailored to specific methodological designs (e.g., randomised controlled trials). These tools thus place great emphasis on assessing methodological features such as participant allocation and blinding that are less standard in human-computer interaction studies. Other popular tools are easy to apply but lack critical assessment criteria, including sample size and study design, such as the Mixed Methods Appraisal Tool (MMAT)83 and Joanna Briggs Institute (JBI) Critical Appraisal Checklists84. Alternatively, some are generalised for mixed- and multi-method studies; examples include the Quality Assessment for Diverse Studies (QuADs)85 and Standard Quality Assessment Criteria for Evaluating Primary Research Papers from a Variety of Fields (QualSyst)86, whose criteria are too broad to allow robust assessment of quantitative studies in a meta-analysis. Therefore, although developing a research quality assessment tool was not the primary goal of this review, we decided to devise a tailored checklist by adapting items from existing popular tools81–86 and drawing on prior work that customised quality assessment tools for their reviews87,88. This 23-item quality checklist is available in Supplementary Table 5.
To further test the impact of research quality on the pooled effect sizes, we accounted for its multidimensional nature rather than collapsing items into a single score89,90. Specifically, we computed three composite quality metrics. First, study design rigour reflects how thoughtfully each study was planned and designed, assessed through quality items on objectives and preregistrations, participants, and study design. Second, data & reporting rigour reflects how rigorously study data were handled, assessed through items on data collection, analysis, and results. Third, broad research integrity captures broader practices that support reproducibility and trustworthiness, assessed through items on discussion, ethics, and open science. The first two metrics directly addressed a study’s risk of bias, while the third related more to the overall research quality.
Data analysis
All analyses were run in R v4.2.2. For each eligible study, we calculated the standardised mean difference between human-agent and human-human interaction conditions for specific human responses, serving as the effect size estimates in subsequent analyses. Hedges’ g was chosen over Cohen’s d to account for the potential bias in estimating effects with small sample sizes. Hedges’ g of 0.1 was interpreted as a negligible effect size, 0.2 as small, 0.5 as medium, and 0.8 as large91. When required statistics for calculating unadjusted effect sizes were missing or contained substantial errors, we contacted the corresponding authors to request clarification and essential statistical details or anonymised data, as per their preference. If authors did not respond or provide the requested information, we excluded those missing effects from the main analyses (and, if a study lacked all effects, we excluded that study in full). Nevertheless, when it was possible to approximate missing effects, either by estimating from reported thresholds (e.g., assuming p = 0.005 for p < 0.005), deriving statistically adjusted estimates, or imputing non-significant effects (i.e., max+, max–, and zero-coded45,57), we retained those approximated effect sizes in sensitivity analysis. Supplementary Table 6 presents all formulae for effect size calculation.
We conducted separate meta-analyses for different types of human responses, proceeding with a specific response type only when five or more studies were included92. We employed random-effects meta-analysis, which assumes that the different studies estimate different yet related effects89. The included studies differed in samples, design, and measures as expected, thus fitting this assumption. The random-effects framework accounts for both sampling error and between-study heterogeneity in effect sizes, yielding a pooled effect size that represents the mean of a distribution of true effects rather than a single fixed effect93. Thus, this framework allows the meta-analytic results to be generalised to a broader “population” of potential studies beyond those included in the analysis. Moreover, to accommodate dependencies between multiple effect sizes per response outcome within studies and avoid imposing the independence assumption, we employed a three-level random-effects model. Three sources of variance were identified: random sampling error (level 1), variance among effect sizes within studies (level 2), and variance among effect sizes between studies (level 3)94. We applied both frequentist and Bayesian approaches in the meta-analysis.
We conducted frequentist meta-analysis via the metafor R package95, employing a restricted maximum likelihood estimator to model heterogeneity and t distribution-based inference. We assessed effect-size heterogeneity using two indicators: (1) Q-statistic, which tests for the presence of heterogeneity via the significance of its p-value, and the (2) I2-statistic, which quantifies the magnitude of the heterogeneity. I2 represents the proportion of total variance among effect sizes attributable to true heterogeneity rather than sampling error, with values of 25%, 50%, and 75% interpreted as low, moderate, and high heterogeneity96. The presence of heterogeneity indicates the need for further moderator analysis (i.e., meta-regression) to explore potential sources of this heterogeneity. Each potential moderator was evaluated individually only when data from at least ten studies were available89. For categorical moderators with multiple conditions, only conditions represented by at least three studies were included97. Meta-regression was also employed to test whether research quality systematically impacted the pooled effect size, with three quality metrics evaluated individually. Subgroup analyses were conducted when meta-regression yielded a significant moderating effect, and were visualised using orchard plots via the orchaRd R package98. Additionally, we checked publication bias for response types with at least ten studies89 by visually inspecting funnel plots for asymmetry and performing Egger Sandwich tests via the clubSandwich R package99, with regression slope significance serving as a statistical indicator of asymmetry. We further conducted sensitivity analysis for outliers (effect sizes with absolute studentised deleted residuals exceeding 1.96100) and influential cases (identified using Cook’s distance, DFBETAS values, and hat values95), re-running all meta-analyses after removing these cases. This evaluated whether results were sensitive to potential deviations from underlying distributional assumptions89, thereby assessing the robustness of pooled effect sizes.
To complement the traditional frequentist meta-analytic approach, we also conducted Bayesian meta-analysis via the brms and bridgesampling R packages101,102. Bayesian analysis allows incorporation of prior information to better estimate different sources of variance and enables direct probability statements about parameters via credible intervals103. We pre-determined prior distributions for effect size estimates and heterogeneity parameters. For effect sizes, we used a Cauchy(0, 1/√2) distribution, which has become a default prior in the field of psychology104, to reflect uncertainty about both the direction and magnitude of effect sizes when comparing responses in human-agent vs. human-human interactions. For between-study variance, we used an inverse-Gamma(1, 0.15) distribution, based on between-study heterogeneity estimates from meta-analyses reported in Psychological Bulletin from 1990–2013104,105. For within-study variance, we used an inverse-Gamma(1, 0.1) distribution, reflecting a priori expectation that variance among effect sizes within studies is smaller than variance between studies. We used a Cauchy(0, 1/√2) distribution for moderator effects. We also assessed the sensitivity of Bayesian meta-analytic results to prior distributions by using the following alternatives. For effect sizes, we considered a Student-t(3, 0, 1) prior, a weakly informative distribution that places less density at zero and more on moderate-to-large effects and has lighter tails than Cauchy(0, 1/√2). We considered half-Cauchy(0, 0.3), a commonly used weakly informative prior106, for between-study variance and half-Cauchy(0, 0.2) for within-study variance. Furthermore, we performed Markov Chain Monte Carlo diagnostics to assess model validity: potential scale reduction statistics for convergence107, and effective sample sizes (ESS ≥ 1,000) for all parameters for sampling efficiency101. In the Bayesian approach, prior distributions constituted explicit modelling assumptions; sensitivity analyses and convergence diagnostics indicated that these assumptions were appropriate for our data.
In addition, the Bayesian approach provides formal measures (i.e., Bayes factors; BF10) of the strength of evidence for the study hypothesis (H1) relative to the null hypothesis (H0), thereby indicating when pooled effects should be interpreted with caution91. Specifically, BF10 < 1 indicates the data are more supportive of H0 than H1, with values of 1/3–1 indicating ambiguous evidence, 1/10–1/3 substantial evidence, 1/30–1/10 strong evidence, 1/100–1/30 very strong evidence, and < 1/100 decisive evidence in support of H0; BF10 = 1 indicates perfect ambiguity; BF10 > 1 indicates the data are more supportive of H1, with values of 1–3, 3–10, 10–30, 30–100, and > 100 indicating ambiguous, substantial, strong, very strong, and decisive evidence in support of H1, respectively108.
Ethics statement
This systematic review and meta-analysis is based on data extracted from previously published studies. Where necessary, additional anonymised data were requested and obtained ethically and legally from the original study authors. No new data were collected from human participants. Thus, no ethical approval was required in accordance with the institutional ethical guidelines.
Results
The study selection process is summarised in the PRISMA flowchart (Fig. 2). After title-abstract screening (11,456 records) and full-text screening (815 articles), 162 studies (from 122 articles) met our eligibility criteria. Key characteristics of eligible studies are provided in Supplementary Table 7. These studies investigated a broad set of psychological and behavioural responses, which we classified into distinct response types and meta-analysed separately. Sixteen studies examined rare responses with insufficient data for meta-analysis92, and these were narratively synthesised in Supplementary Table 8. The final quantitative synthesis included 146 studies (from 112 articles), yielding 468 effect sizes (3.21 per study).
Fig. 2. PRISMA flowchart.
Two articles165,166 reported the same underlying study while analysing different human responses, and were treated as one study in our meta-analysis. Articles167,168 likewise reported the same study and were treated as one. There was one article169 describing a single investigation but collecting and analysing data separately for Chinese and US samples; to keep sample independence, we treated these as two studies. Another article170 reported Japanese and US samples, which we also treated as two studies. In addition, of the 25 articles excluded for underreported or inappropriate statistical information, 16 articles (22 studies) allowed deriving approximated effect sizes for some or all responses examined. These approximated effect sizes were included in the sensitivity analyses in the supplementary information.
Integrating frequentist and Bayesian approaches, we conducted a series of meta-analyses comparing individuals’ psychological and behavioural responses in dyadic interactions with performance-matched agent vs. human partners. Studies included in the meta-analyses were published between 2003 and 2024 across 112 publications across diverse outlets. Most publications (85) were journal articles, with Computers in Human Behavior (10) and International Journal of Social Robotics (8) among the more frequently represented outlets. The remaining 27 appeared in conference proceedings, with the ACM International Conference on Intelligent Virtual Agents (IVA; 4) and the IEEE International Conference on Robot and Human Interactive Communication (RO-MAN; 4) being relatively common venues.
Of 146 included studies, the average sample size was 162.80 (SD = 163.13, median = 116.17, range = 8–945). For seven studies that reported only total sample sizes and included conditions beyond human-agent and human-human interactions, we assumed equal condition sizes109,110; removing these yielded a similar average of 164.60 (SD = 166.21, median = 117). Using sample-size weighting, the mean participant age was 31.62 years (range = 18.60–78.79), and the mean percentage of female participants was 56.87% (range = 0–100%). The study sample was geographically diverse, while Western countries and East Asia were disproportionately represented. Most studies were conducted in the USA (50) and China (18), followed by Japan (12), Germany (8), Italy (8), and the UK (5). Other locations included France (2), Finland (2), and several countries represented by one study each (Australia, Belgium, Canada, New Zealand, Singapore, South Korea, Sweden, and the Netherlands). Six studies involved multi-country samples, and 27 studies did not report location information.
RQ1: Which psychological and behavioural responses have been investigated in studies comparing human-agent and human-human interactions?
We identified 23 types of human responses, each with at least five studies for meta-analysis92. They were grouped into six themes: prosociality and morality, social perceptions of interaction partners, trust in interaction partners, social alignment with interaction partners, personal agency and task performance, and interaction experiences. An overview of the response types, grouped under these themes and with detailed conceptualisations, is provided in Table 2.
Table 2.
Overview of the meta-analysed response types
| Response theme | Response | Conceptualised As |
|---|---|---|
| Prosociality and morality | Prosocial behaviour | Voluntary actions that are intended to help or benefit others111. |
| Moral engagement | Psychological and behavioural commitment to moral standards in a given context, demonstrating direct or indirect responsiveness to the needs and interests of others112,113. | |
| Social perceptions of interaction partners | Perceived social presence | Subjective experience of being present with a real social partner, with whom one can exchange thoughts and emotions56,153. |
| Perceived likeability | Positive evaluation of the partner based on their affiliative capacity and social attractiveness115,116. | |
| Perceived competence | Evaluation of the partner’s intelligence and domain-specific knowledge, or their effectiveness and efficiency in task performance115,117. | |
| Agency attribution | Explicit, reflective process of ascribing the partner the capacity to intentionally initiate events, and recognising them as the cause of their own decisions or actions118. | |
| Responsibility attribution | Process of assigning responsibility, with credit or blame, for an event or action to the partner119. It is closely related to agency attribution, which lays the foundation: the partner needs to be perceived as having agency to cause an outcome before being assigned responsibility for it154,155. | |
| Trust in interaction partners | Behavioural trust | Observable actions of voluntarily accepting vulnerability based on positive expectations of the partner, such as risk-taking behaviour or reliance on the partner under uncertainty156,157. |
| Subjective trust | Psychological state comprising the intention to accept vulnerability based on positive expectations of the partner, such as their ability, benevolence, or integrity120,157. | |
| Social alignment with interaction partners | Social alignment | Voluntary alignment of internal states and behaviours with the partner, including self-other integration (merging identities and perspectives), advice taking (incorporating partner insights), synchrony (adapting verbal and nonverbal behaviours), and proximity regulation (regulating interpersonal distances)121. |
| Personal agency and task performance | Perceived self-agency | Subjective experience of controlling one’s own body and external events, which leads one to feel responsible for what their decisions or actions cause158. |
| Self-disclosure | Process of revealing personal information, such as one’s own thoughts, feelings, and experiences, with the partner159. | |
| Strategic economic behaviour | Deliberate choices of actions in economic games where one, recognising the interdependence of actions, anticipates and reacts to the partner’s actions by weighing how their choices would affect both personal and partner outcomes. It can be affected by multiple factors, including not only payoff maximisation, but also social preferences (e.g., reciprocity, trust, and fairness) or even intuitions123,160. | |
| Objective task performance | Measurable and quantifiable assessments regarding one’s performance of specific tasks. | |
| Interaction experiences | Perceived partner relational qualities | Subjective evaluations of the relational attributes the partner exhibits during interaction, including perceived rapport (closeness and mutual connection), perceived interactivity (active engagement and responsiveness), perceived empathy and supportiveness (understanding of and support for one’s perspectives and feelings), and perceived customer orientation (commitment to meeting one’s needs)124. |
| Affective valence | Hedonic tone (i.e., unpleasantness–pleasantness) of the emotional experience when interacting with the partner161. | |
| Affective arousal | Activation level (i.e., calmness–excitement) of the emotional experience when interacting with the partner161. | |
| Interaction satisfaction | Overall evaluation of how well the interaction with the partner meets one’s needs and expectations125. | |
| Future interaction intention | Degree to which one would like to interact with the partner again in the future162. | |
| Perceived interaction naturalness | Degree to which one perceives their interaction with the partner as natural. | |
| Perceived interaction enjoyment | Degree to which one feels they have enjoyed interaction with the partner. | |
| Subjective workload | Perception and emotional experience of the overall effort invested in performing specific tasks163. | |
| Subjective task engagement | Perception and emotional experience of involvement and investment in performing specific tasks164. |
To derive response types and themes, we performed a posteriori classification of the diverse human responses investigated across studies. After a conceptual-to-empirical approach79 to classification development failed to align with extracted responses (Supplementary Note 1), we adopted an empirical-to-conceptual approach79, allowing the classification scheme to arise directly from the dataset. We reviewed the descriptions and measures of all human responses extracted from the studies and inductively classified them into distinct response types. Conceptually aligned response types were further grouped into six identified themes, with a residual category for unclassified responses.
RQ2: Which specific response types differ between human-agent and human-human interactions?
Meta-analyses revealed notable differences in prosociality and morality, as well as in social perceptions of partners, between human-agent and human-human interactions. Detailed results are summarised in Table 3, with pooled effect sizes visualised in Fig. 3.
Table 3.
Meta-analytic results for different response types in human-agent vs. human-human interactions
| Response theme | Response | Hedges’ g | 95% CI | t | t_p | Q | Q_p | I2 (%) | BF10 | k | m |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Prosociality and morality | Prosocial behaviour | −0.648 | [−0.82, −0.48] | −8.66 | < 0.001 | 10.83 | 0.288 | 21.78 | 6003.065 | 9 | 10 |
| Moral engagement | −0.376 | [−0.48, −0.27] | −8.19 | < 0.001 | 4.04 | 0.991 | 0 | 1725.850 | 8 | 14 | |
| Social perceptions of interaction partners | Perceived social presence | −0.284 | [−0.52, −0.05] | −2.81 | 0.023 | 48.72 | < 0.001 | 68.19 | 3.495 | 9 | 20 |
| Perceived likeability | −0.352 | [−0.54, −0.16] | −3.78 | < 0.001 | 443.07 | < 0.001 | 88.39 | 41.417 | 28 | 41 | |
| Perceived competence | −0.457 | [−0.61, −0.30] | −6.21 | < 0.001 | 148.84 | < 0.001 | 77.56 | 10296.030 | 23 | 31 | |
| Agency attribution | −0.705 | [−1.16, −0.25] | −3.43 | 0.006 | 229.52 | < 0.001 | 94.62 | 16.554 | 11 | 18 | |
| Responsibility attribution | −0.511 | [−0.72, −0.30] | −5.44 | < 0.001 | 92.77 | < 0.001 | 81.56 | 512.117 | 12 | 18 | |
| Trust in interaction partners | Behavioural trust | −0.020 | [−0.16, 0.12] | −0.30 | 0.769 | 47.98 | 0.002 | 60.83 | 0.077 | 18 | 24 |
| Subjective trust | −0.109 | [−0.23, 0.01] | −1.88 | 0.073 | 115.41 | < 0.001 | 76.87 | 0.358 | 25 | 35 | |
| Social alignment | Social alignment | −0.004 | [−0.10, 0.09] | −0.08 | 0.937 | 52.90 | 0.027 | 33.11 | 0.052 | 22 | 36 |
| Personal agency and task performance | Perceived self-agency | 0.010 | [−0.14, 0.16] | 0.16 | 0.880 | 18.51 | 0.139 | 37.53 | 0.084 | 9 | 14 |
| Self-disclosure | 0.025 | [−0.16, 0.21] | 0.31 | 0.766 | 28.03 | 0.109 | 44.89 | 0.089 | 9 | 21 | |
| Strategic economic behaviour | −0.086 | [−0.18, 0.01] | −1.93 | 0.078 | 32.82 | 0.084 | 32.26 | 0.246 | 13 | 24 | |
| Objective task performance | 0.008 | [−0.10, 0.12] | 0.15 | 0.884 | 70.60 | < 0.001 | 46.39 | 0.062 | 23 | 38 | |
| Interaction experiences | Perceived partner relational qualities | −0.157 | [−0.40, 0.09] | −1.36 | 0.194 | 159.82 | < 0.001 | 91.66 | 0.320 | 17 | 25 |
| Affective valence | −0.226 | [−0.57, 0.12] | −1.49 | 0.171 | 84.45 | < 0.001 | 86.33 | 0.410 | 10 | 14 | |
| Affective arousal | −0.087 | [−0.34, 0.17] | −0.87 | 0.422 | 17.81 | 0.058 | 46.56 | 0.168 | 6 | 11 | |
| Interaction satisfaction | 0.085 | [−0.12, 0.29] | 0.86 | 0.400 | 281.59 | < 0.001 | 94.73 | 0.154 | 19 | 27 | |
| Future interaction intention | 0.105 | [−0.10, 0.31] | 1.32 | 0.245 | 14.10 | 0.028 | 57.34 | 0.212 | 6 | 7 | |
| Perceived interaction naturalness | −0.368 | [−1.24, 0.50] | −1.18 | 0.304 | 73.40 | < 0.001 | 96.97 | 0.585 | 5 | 5 | |
| Perceived interaction enjoyment | −0.182 | [−0.69, 0.32] | −0.88 | 0.411 | 83.97 | < 0.001 | 89.06 | 0.355 | 7 | 10 | |
| Subjective workload | −0.157 | [−0.54, 0.23] | −1.13 | 0.323 | 32.49 | 0.002 | 69.04 | 0.320 | 5 | 14 | |
| Subjective task engagement | −0.063 | [−0.31, 0.19] | −0.57 | 0.585 | 20.85 | 0.022 | 58.66 | 0.168 | 10 | 11 |
For simplicity, this table presents results primarily from frequentist meta-analysis. BF10 is presented to provide complementary Bayesian evidence for H1 over H0. All frequentist estimates are consistent with Bayesian estimates. Full Bayesian results are presented in Supplementary Table 16. Specifically, Hedges’ g, 95% CI, t-value and associated p-value were estimated via random-effects meta-analysis. Q is Cochrane’s Q-statistic for testing heterogeneity. I2 is the proportion of total variance attributable to true heterogeneity rather than sampling error. k is the number of studies including in meta-analysis. m is the number of effect sizes included.
Fig. 3. Forest plot visualising pooled effect sizes for different response types.
Only response types with five or more studies were included in the meta-analysis and are shown in this plot. Horizontal bars represent 95% CIs. Response types with bars that do not cross the dashed vertical “line of no effect” indicate existing significant differences between human-agent and human-human interactions. Ticks on the x-axis denote conventional thresholds for small (g = 0.2), medium (g = 0.5), and large (g = 0.8) effects. The pooled effect sizes shown in this plot were derived from frequentist meta-analyses, with the accompanying strength of Bayesian evidence indicated by coloured circles. To aid interpretation, effect sizes were also pooled at the theme level and shown as diamonds. For the response theme comprising one response type (i.e., social alignment), the response-level estimate also represents the theme-level pooled effect. Full theme-level meta-analytic results are presented in Supplementary Table 17.
Prosociality and morality
Prosocial behaviour refers to voluntary actions intended to benefit others111, with frequentist meta-analysis revealing a medium-to-large partner effect. Individuals behaved significantly less prosocially when interacting with agent vs. human partners, for example by sharing less in dictator games or adapting less to the partner’s perspective in joint tasks. Moral engagement, the psychological and behavioural commitment to moral standards112,113, showed a small-to-medium partner effect. Individuals engaged significantly less morally with agent vs. human partners, with reduced moral acts, intentions, or feelings of guilt.
Social perceptions of interaction partners
On average, perceived social presence—the sense of the partner being “there” and socially real—showed a small partner effect, with agents perceived as significantly less socially present than human partners. For two core dimensions of social perception114, perceived likeability reflects affective evaluations of the partner115,116 (e.g., warmth, friendliness), while perceived competence captures evaluations of capability115,117 (e.g., intelligence, effectiveness). Both showed small-to-medium partner effects, with agents perceived as significantly less likeable and competent than human partners. The agent disadvantage was greater for higher-order social constructs: a medium-to-large partner effect on agency attribution—capacity for intentional action118—and a medium effect on responsibility attribution—accountability for outcomes119, with agents attributed significantly less agency and responsibility.
Bayesian meta-analytic results were consistent with frequentist estimates. Bayesian support was decisive for partner effects on prosocial behaviour and moral engagement. Regarding social perceptions, support was substantial for partner effects on perceived social presence, strong for agency attribution, very strong for perceived likeability, and decisive for perceived competence and responsibility attribution.
RQ3: Which specific response types are similar between human-agent and human-human interactions?
Meta-analyses revealed general response similarities between human-agent and human-human interactions across four response themes: trust, social alignment, perceived agency and task performance, and interaction experiences. Detailed results are provided in Table 3 and Fig. 3.
Trust in interaction partners
Trust is “a willingness to be vulnerable” in the absence of the ability to monitor the trustee120, encompassing both behavioural and psychological aspects. Frequentist meta-analyses revealed no significant differences in behavioural trust (e.g., risk-taking or reliance on the partner under uncertainty) and subjective trust (e.g., perceived trustworthiness or reliability) towards agent vs. human partners.
Social alignment with interaction partners
Social alignment refers to the voluntary alignment of internal states and behaviours with the partner121, such as self-other integration and behavioural synchrony. Like behavioural trust, it is critical for interaction coordination121 and showed no significant difference with agent vs. human partners.
Personal agency and task performance
Perceived self-agency and self-disclosure are self-oriented processes122 during interaction. Self-agency is the sense of control over one’s own body and external events, while self-disclosure involves sharing personal information. Both showed no significant differences during agent vs. human interactions. Moreover, strategic economic behaviour—deliberate choices in economic games while acknowledging the interdependence of players’ actions123—did not differ significantly by partner type. Objective task performance—quantifiable assessments of task execution (e.g., accuracy or response time)—also did not differ significantly.
Interaction experiences
On average, perceived relational qualities of agent vs. human partners—subjective evaluations of relational attributes the partner exhibits during interaction124, such as rapport and empathy—did not differ significantly. Likewise, no significant differences were found for affective valence (e.g., unpleasantness–pleasantness) and arousal (e.g., calmness–excitement). Interaction satisfaction—overall evaluation of how well the interaction meets one’s needs and expectations125—did not differ significantly by partner type, nor did future interaction intention (e.g., willingness to engage again), perceived naturalness, or enjoyment of interaction. For subjective task experience, subjective workload (e.g., perceived effort) and task engagement (e.g., enjoyment or immersion in the task) did not differ significantly.
Bayesian meta-analytic results were consistent with frequentist estimates. However, Bayesian evidence was ambiguous for the absence of partner effects on subjective trust, affective valence, perceived interaction naturalness, and enjoyment. For other response types, evidence for the absence of partner effects was clearer. Bayesian support was strong for behavioural trust, social alignment, perceived self-agency, self-disclosure, and objective task performance, and substantial for strategic economic behaviour and most interaction experiences, including perceived partner relational qualities, interaction satisfaction, future interaction intention, affective arousal, subjective workload, and task engagement.
RQ4: To what extent do study, participant, partner, interaction, and response characteristics moderate partner effects on different response types?
We conducted univariate meta-regressions exploring characteristics of studies, participants, partners, interactions, and responses as moderators (see Methods for moderator coding). Of 23 response types meta-analysed, moderator analyses were not conducted for partner effects on prosocial behaviour, moral engagement, perceived self-agency, self-disclosure, strategic economic behaviour, and affective arousal due to non-significant heterogeneity in effect sizes (Q-test p < 0.05). Despite significant effect-size heterogeneity, the limited number of studies (k < 10)89 precluded moderator analyses for perceived social presence, future interaction intention, perceived interaction naturalness and enjoyment, and subjective workload. Therefore, we conducted moderator analyses for 12 response types: perceived likeability, perceived competence, agency attribution, responsibility attribution, behavioural trust, subjective trust, social alignment, objective task performance, perceived partner relational qualities, affective valence, interaction satisfaction, and subjective task engagement.
For simplicity, this section only presents results for six response types under the themes of “social perceptions of interaction partners” and “interaction experiences” (Tables 4 and 5), with moderators both reaching frequentist significance and receiving at least substantial Bayesian support (p < 0.05; BF10 > 3). Full results are provided in Supplementary Table 9.
Table 4.
Meta-regression and subgroup analysis results for social perceptions of interaction partners
| Response | Meta-regression | Subgroup analysis |
|---|---|---|
| Perceived likeability | Appearance difference: F = 7.48, p = 0.011; BF10 = 4.559 | ![]() |
| Interaction task: F = 5.16, p = 0.007; BF10 = 4.839 | ![]() |
|
| Perceived competence | Agent form: F = 13.42, p = 0.001; BF10 = 16.598 | ![]() |
| Appearance difference: F = 12.49, p = 0.002; BF10 = 19.646 | ![]() |
|
| Interaction medium: F = 12.08, p = 0.002; BF10 = 14.841 | ![]() |
|
| Agency attribution | Interaction medium: F = 7.09, p = 0.026; BF10 = 4.331 | ![]() |
| Responsibility attribution | Study setting: F = 11.34, p = 0.007; BF10 = 4.271 | ![]() |
| Interaction realism: F = 11.34, p = 0.007; BF10 = 4.271 | ![]() |
Moderator analyses were conducted for response types with significant effect-size heterogeneity and at least ten available studies. Potential moderators were tested individually. For categorical moderators with multiple conditions, only conditions represented by at least three studies were included. For simplicity, this table presents results only for moderators that both reached frequentist significance and received at least substantial Bayesian support (p < 0.05; BF10 > 3). Full results for all tested moderators are available in Supplementary Table 9. For meta-regression results, F-value is from random-effects meta-regression. BF10 represents the Bayesian evidence for H1 (i.e., presence of a moderating effect) over H0 (i.e., absence of a moderating effect). For follow-up subgroup analyses, Hedges’ g and 95% CI were estimated via random-effects meta-analyses and visualised via orchard plots. k is the number of studies including in the analysis. m is the number of effect sizes included.
Table 5.
Meta-regression and subgroup analysis results for interaction experiences
| Response | Meta-regression | Subgroup analysis |
|---|---|---|
| Perceived partner relational qualities | Agent operationalisation:F = 11.37, p = 0.001; BF10 = 71.227 | ![]() |
| Human partner type:F = 10.75, p = 0.002; BF10 = 45.469 | ![]() |
|
| Interaction realism:F = 23.50, p < 0.001; BF10 = 298.645 | ![]() |
|
| Response dimension:F = 6.78, p = 0.006; BF10 = 3.334 | ![]() |
|
| Interaction satisfaction | Interaction nature:F = 40.31, p < 0.001;BF10 = 15137.576 | ![]() |
The same note as in Table 4 applies.
Social perceptions of interaction partners
For perceived likeability, appearance difference (matched vs. differed) and interaction task (service encounter vs. game vs. instructional interaction vs. communication-focused) moderated the partner effect. Agents were perceived as significantly less likeable than human partners (a medium partner effect) when appearance differed, but this was non-significant when appearances matched. Moreover, agents were significantly less likeable than human partners in service encounters and games (medium and high partner effects, respectively), but this was non-significant in instructional interactions and communication-focused interactions.
For perceived competence, appearance difference (matched vs. differed), agent form (physical vs. virtual), and interaction medium (face-to-face vs. computer-mediated) moderated the partner effect. Agents were perceived as significantly less competent than human partners (a medium-to-large partner effect) when appearance differed, but this was non-significant when appearances matched. The agent disadvantage manifested across conditions of agent form and interaction medium but was significantly greater for physical agents than virtual agents (medium-to-large vs. small) and greater in face-to-face than computer-mediated interactions (medium-to-large vs. small).
For agency attribution, interaction medium (face-to-face vs. computer-mediated) moderated the partner effect. Agents were attributed significantly less agency than human partners (a large partner effect) in face-to-face interactions; this partner effect was small-to-medium while non-significant in computer-mediated interactions.
For responsibility attribution, study setting (online vs. lab) and interaction realism (hypothetical vs. real-time)—which were perfectly co-linear—moderated the partner effect. In online studies with hypothetical interaction, agents were attributed significantly less responsibility than human partners (a medium-to-large partner effect), whereas lab studies with real-time interaction yielded a non-significant partner effect.
Interaction experiences
For perceived partner relational qualities, agent operationalisation (vignette-described vs. Wizard-of-Oz vs. autonomous), human partner type (vignette-described partner vs. research team member vs. algorithm-controlled pseudo-human), and interaction realism (hypothetical vs. real-time) moderated the partner effect. Perfect multicollinearity occurred among moderator conditions due to a subset of studies that asked participants to imagine interacting with partners (i.e., hypothetical interaction) through vignettes. This subset yielded a medium-to-large partner effect, reflecting methodological contribution to agent disadvantage. Studies with real-time interaction yielded a non-significant partner effect, regardless of agent operationalisation or human partner type. Additionally, response dimension moderated this partner effect. Studies measuring perceived customer orientation showed that agents were perceived as significantly less customer-oriented than human partners (a medium-to-large partner effect), whereas studies of other relational qualities—perceived rapport, interactivity, and empathy—yielded non-significant partner effects.
For interaction satisfaction, interaction nature (oppositional vs. cooperative) moderated the partner effect. In oppositional interactions, satisfaction was significantly higher with agent vs. human partners (a medium partner effect), whereas in cooperative interactions, this was non-significant.
In brief, no participant characteristics demonstrated robust moderating effects (i.e., effects reaching frequentist significance with at least substantial Bayesian support). Nevertheless, significant moderators of partner effects on specific response types were identified across other categories: study characteristics (study setting), response characteristics (response dimension), partner characteristics (appearance difference, agent form, agent operationalisation, and human partner type), as well as interaction characteristics (interaction task, medium, realism, and nature).
Publication bias, research quality, and sensitivity analysis
For response types with at least ten studies89, we evaluated publication bias via funnel plots and Egger Sandwich tests126. Visual inspection of funnel plots showed fairly symmetric distributions for most response types (Supplementary Fig. 1), which were also supported by non-significant Egger tests (Supplementary Table 10). Thus, the risk of publication bias was minimal. There were two exceptions: social alignment (Egger p = 0.018, with its funnel showing a missing left wedge and right-skewed clustering of smaller studies) and objective task performance (Egger p = 0.038, despite a relatively symmetric funnel). Nevertheless, given near-zero pooled effect sizes and mostly non-significant individual effects for both response types, the observed asymmetry likely reflected small-study effects (i.e., small samples are associated with greater random error and yield larger yet mostly non-significant effect estimates with wide confidence intervals, resulting in apparent asymmetry) rather than true publication bias. Research quality was assessed via a tailored checklist, yielding three quality metrics—study design rigour, data & reporting rigour, and broad research integrity—for each study (Supplementary Table 11). Meta-regressions showed no systematic impact of research quality on our results, though modest influences of specific quality metrics cannot be ruled out; details are provided in Supplementary Table 12. Finally, three sensitivity checks—incorporating approximated effect sizes, removing outliers and influential cases, and using alternative Bayesian priors—confirmed that our main meta-analytic results were robust across analytic decisions; details are provided in Supplementary Tables 13–15.
Discussion
This review synthesised empirical evidence on similarities and differences in individuals’ psychological and behavioural responses when interacting with performance-matched agent vs. human partners. We conducted separate random-effects meta-analyses for 23 response types, followed by univariate meta-regressions for 12 response types with significant effect-size heterogeneity and at least ten available studies.
RQ1: Which psychological and behavioural responses have been investigated in studies comparing human-agent and human-human interactions?
In total, we identified 23 types of psychological and behavioural responses across six themes, reflecting the breadth of human responses examined in human-agent and human-human interaction research. However, the distribution of response types across studies was uneven. Perceived likeability and competence, subjective trust, social alignment, and objective task performance were among the most frequently investigated response types. In contrast, certain interaction experiences, including affective arousal, future interaction intention, perceived interaction naturalness and enjoyment, and subjective workload, were compared less frequently in human-agent vs. human-human interactions.
Responses within the themes of “social perceptions of interaction partners” and “interaction experiences” were particularly diverse, with many lacking sufficient data for meta-analysis. They were measured using heterogeneous instruments that captured different aspects of constructs at varying levels of abstraction, contributing to wider confidence intervals in pooled effect sizes. For example, interaction satisfaction has been assessed with measures ranging from single- to multi-item scales, raising concerns about their psychometric comparability. This aligns with a previous review that identified a continued tendency to develop bespoke questionnaires rather than adopt validated ones127. Notably, less empirical research has examined “negative” human responses, despite large-scale self-report evidence showing a range of negative responses to agents23. Greater attention to responses such as hostility, blame, and distrust is needed, as real-world interactions are not uniformly positive, and these reactions may be distinct from, rather than the inverse of, positive responses.
RQ2: Which specific response types differ between human-agent and human-human interactions?
Our meta-analyses found that individuals were less prosocial and morally engaged when interacting with agents than with human partners. These results quantitatively substantiate and extend a systematic review showing increased selfishness and rationality when playing behavioural economic games with computer vs. human partners128, suggesting that reduced prosocial behaviour and moral engagement with agents generalises beyond economic paradigms to broader interaction contexts.
Compared to human partners, agents were on average perceived as possessing fewer social attributes: they were rated as less likeable and competent, and attributed less agency and responsibility. Agents were also perceived as having lower social presence. However, this partner effect was small and sensitive to how missing non-significant effects were handled, necessitating further research to reach robust conclusions regarding perceived social presence. This pattern aligns with prior meta-analytic findings that agent anthropomorphism was more strongly associated with user perceptions such as likeability and intelligence than with social presence72. Overall, our review provides quantitative support for theoretical accounts that current intelligent agents lack key social affordances necessary to cultivate fully human-like relationships27.
RQ3: Which specific response types are similar between human-agent and human-human interactions?
Alongside response differences, we identified key response similarities between human-agent and human-human interactions. Specifically, individuals demonstrated social alignment and behavioural trust with agent and human partners comparably. Subjective trust likewise showed no significant difference in frequentist meta-analysis, whereas weak Bayesian support for the null effect leaves open the possibility of a modest partner effect. Previous quantitative syntheses demonstrated positive effects of anthropomorphism on trust attitudes towards agents73,74. However, these syntheses included many studies in which participants evaluated agents without interacting with them. Trust formed in evaluative paradigms may reflect initial impressions rather than trust perceptions and behaviours that develop through interaction. Our work extends this literature by demonstrating that trust does not reliably differ between two interactions, consistent with a prior review proposing that trust in agent and human partners develops through similar underlying processes59. In addition, individuals exhibited comparable self-agency, self-disclosure, and strategic economic behaviour in both interactions. Objective task performance was likewise comparable, in line with a meta-analysis of agent anthropomorphism showing no significant effect on participants’ task performance74.
Most interaction experiences were on average similar between partner types. Bayesian support for no partner effects was ambiguous for affective valence and perceived interaction naturalness and enjoyment. Therefore, modest partner effects for these three responses cannot be ruled out. Echoing this uncertainty, prior meta-analyses regarding affective valence have shown mixed findings: one synthesis reported a significant effect of agent anthropomorphism on overall affective valence74, while others found significant associations with positive affect but not negative affect72,73.
RQ4: To what extent do study, participant, partner, interaction, and response characteristics moderate partner effects on different response types?
Many subjective responses (i.e., social perceptions of partners, subjective trust, and interaction experiences) exhibited high effect-size heterogeneity, indicating that the average effects from meta-analyses should not be interpreted as universal, with true partner effects being context-dependent. Across examined social attributes, agents were on average perceived less positively than human partners, but several moderators shaped these partner effects. Specifically, the partner effect on perceived likeability disappeared when partner appearance matched (vs. differed) or in instructional/communication-focused interaction (vs. game/service encounter) tasks. The partner effect on perceived competence disappeared when partner appearance matched (vs. differed) and was reduced when agents were virtual (vs. physical) or in computer-mediated (vs. face-to-face) interactions. Meta-analyses on agent anthropomorphism and social cues also reported positive associations with user perceptions, including likeability and intelligence72–74, suggesting that these agent disadvantages relate to design cues rather than inherent agent limitations. In addition, partner effects on agency and responsibility attribution disappeared in computer-mediated (vs. face-to-face) interactions and in lab-setting, real-time (vs. online-setting, hypothetical) interactions, respectively.
Given the limited studies for many interaction experiences, moderators were identified only for interaction satisfaction and perceived partner relational qualities. Interaction satisfaction was higher with agents than with human partners in oppositional (vs. cooperative) interactions. Agents were perceived as having lower relational qualities than human partners in hypothetical, vignette-described (vs. real-time) interactions, or when focusing on customer orientation (vs. interactivity/rapport/empathy). Of note, rapport did not differ significantly between agent and human partners in our review. This adds to existing mixed meta-analytic evidence, with one synthesis finding no significant association between anthropomorphism and rapport with agents72, while another finding a significant effect of human-like social cues on rapport73. For many other tested moderators, evidence regarding their influence was inconclusive—for example, some reached frequentist significance but with weak Bayesian support (e.g., measurement timing for subjective trust). These inclusive results highlight the need for further research to clarify variations in these subjective responses. Overall, we believe the partner effects on these responses should be interpreted by considering both the average effects and the strong effect-size heterogeneity.
Limitations
These findings should be interpreted considering the review limitations. First, despite all studies being peer-reviewed, research quality varied greatly, with three quality metrics averaging only moderate levels. While we found no systematic impact of research quality on results, missing or inconsistent statistics and methodological details in some studies could add noise to certain meta-analytic estimates. Second, our response classification was derived inductively from the variables and measures reported in the included studies and thus cannot capture the full spectrum of human responses possible in interactions with agent and/or human partners. Third, we restricted moderator analyses to univariate models. Multivariate models, even Bayesian regularised meta-regression129, would be uncertain due to a lack of sufficient studies to accommodate multiple moderators simultaneously. Additionally, many subjective responses exhibited high effect-size heterogeneity, which our meta-regressions only partly explained. Measurements used across studies varied in format, reliability, and validity, but they were too diverse to be fully captured in our moderator coding, potentially obscuring nuanced moderating effects.
Fourth, our findings are limited by the scope of this review. We focused on responses in dyadic interactions, so generalising our partner-effect findings to multi-party scenarios requires caution. We analysed individual-level responses directly related to the interaction, whereas dyad-level responses and those reflecting downstream outcomes were beyond the scope. Moreover, most included studies investigated one-time interactions in controlled lab or online settings; responses in sustained, naturalistic interactions remain unclear. Future work should revisit partner effects using more sophisticated agents capable of long-term interaction. This was partly due to earlier technological constraints, with agents in most included studies relying on traditional machine learning, rule-based algorithms, or Wizard-of-Oz setups. Although our inclusion of “intelligent agent” was technology-agnostic, ranging from traditional technologies to generative AI, none of the studies based on generative AI that we initially identified met our eligibility criteria. Nonetheless, we expect that the growing prevalence of generative AI will spur new comparisons of interactions with generative AI agents vs. humans, enabling re-examination and extension of the current findings. In addition, our review covered only general adult populations and was dominated by WEIRD samples130, with limited non-WEIRD representation mainly from East Asia. Future reviews should target younger populations and include broader non-WEIRD populations by synthesising non-English studies.
Methodological strengths
Despite its limitations, this review demonstrates several methodological strengths. It is a comprehensive meta-analysis comparing a broad range of individuals’ psychological and behavioural responses when interacting with performance-matched agent vs. human partners. We rigorously classified the diverse responses using an empirical-to-conceptual approach79. We initially sought to apply a conceptual-to-empirical approach by mapping extracted responses onto an existing taxonomy of psychological and behavioural responses to robots. However, this top-down strategy did not adequately capture the diversity and specificity of responses identified from the included studies. We therefore shifted to an empirical-to-conceptual approach, allowing the classification scheme to emerge from the data. This methodological transition reflects the evolving and interdisciplinary nature of human-agent interaction research. While such agility can introduce uncertainty, particularly in synthesis work, we believe the data-driven approach enabled a more faithful representation of the responses studied and better supported subsequent meta-analytic decisions.
In addition, alongside frequentist analyses, we conducted Bayesian meta-analyses to confirm result consistency. The Bayesian approach also provided complementary insights: Bayes factors quantified the strength of evidence for effect presence/absence, helping distinguish null effects from underpowered results and flag ambiguous cases where pooled effects warrant caution. Sensitivity checks—excluding outliers and influential cases, incorporating approximated effect sizes, and testing alternative Bayesian priors—confirmed the robustness of our main results across analytic decisions. We further tested many potential moderators to explore sources of significant effect-size heterogeneity. Finally, recognising that existing research appraisal tools were less suited to our review, we developed a tailored checklist to assess research quality while capturing its multidimensional nature. The checklist covered quality criteria related to objectives and preregistration, participants, study design, data collection and analysis, results, discussion, as well as ethics and open science. However, certain procedural indicators of research validity (e.g., manipulation checks, inter-rater reliability, and research team consensus procedures) were not incorporated as standalone assessment items. These indicators were not consistently applicable across the included studies or were implemented in varied ways with insufficient reporting detail, limiting their suitability for categorical quality assessment in the current review. Nevertheless, assessment of research validity would likely benefit from considering these indicators, which future reviews may incorporate as reporting practices become more standardised.
Implications
This review has significant theoretical implications. By synthesising empirical evidence across human-computer interaction, social robotics, psychology, communication, and business studies, we established a cross-disciplinary, systematic understanding of similarities and differences in human responses between human-agent and human-human interactions. Our results confirmed that social responses should be considered as distinct constructs rather than a monolithic entity61. Among these social responses, behavioural trust and social alignment showed convergence across interactions, whereas social perceptions, prosociality, and morality diverged. These findings provide partial support for both the Media Equation49 and the Threshold Model of Social Influence51 and refine their scope by demonstrating that social equivalence varies across response types. In addition, the Threshold Model distinguishes between automatic and deliberate processing underlying human responses51, but cognitive processing likely shifts dynamically between these modes during interaction. Although outside the scope of this review, theorising may be expanded to incorporate the temporal nature of human responses.
The review also provides insight into cue-based frameworks, such as the Modality-Agency-Interactivity-Navigability (MAIN) model55 and the Media Are Social Actors (MASA) paradigm75, by identifying empirical boundary conditions for partner effects in interaction. Specific partner and interaction characteristics moderated partner effects on social perceptions and interaction experiences, while showing limited moderating influence on other response types. Our review suggests that theorising the influence of social cues may benefit from differentiating response types and attending to research paradigms (e.g., interaction-based or evaluative). Moreover, by spanning diverse agent morphologies (e.g., robots, virtual humans, conversational agents), we found that partner effects on most responses, such as trust and social alignment, were consistent regardless of agent embodiment (disembodied vs. embodied) or form (physical vs. virtual). This indicates that trust and social alignment, as core collaborative mechanisms121,131, may operate in a morphology-agnostic manner. Prior meta-analyses provide mixed but broadly convergent evidence. Whether robots were embodied or depicted did not significantly moderate the effect of anthropomorphism on attitudinal outcomes, including trust74. Embodiment was also not found to moderate cue-related effects on trust, but physical presence did57. By contrast, a meta-analysis directly comparing physically co-located and virtually represented robots found no significant difference in trust132. Therefore, trust-related responses tend to be relatively robust across agent embodiment and form.
Furthermore, our review benchmarks human responses in human-agent vs. human-human interactions prior to generative AI. Looking ahead, as agents powered by large language models (LLMs) become increasingly integrated into everyday life, sustained and personalised interactions with these agents may attenuate perceived deficits in social attributes. However, this convergence may not extend to prosociality and morality. Reduced prosocial behaviour and moral engagement in human-agent interaction may reflect judgements about moral standing and intrinsic value—issues linked to an entity’s ontological identity133. As agents become more humanlike and autonomous, questions surrounding identity could become more, rather than less, salient. For example, research has found that although LLM-generated messages make participants feel emotionally supported, this support is devalued once the source is identified as AI rather than human134. We also found that functional behaviours and interaction experiences with agents resemble those with humans. Generative AI may not only maintain this parity but, in certain contexts, shift it. These agents deploy superhuman capabilities while interacting without evoking ego threat or social evaluation concerns135, potentially increasing users’ willingness to trust and align with them. A recent study shows that participants can distinguish between LLM- and lawyer-generated legal advice yet still prefer the LLM’s advice136, indicating that trust in AI agents may exceed human-human baselines in some domains. Overall, our review provides a benchmark for future research on interaction with generative AI agents, enabling researchers to detect potential shifts in human psychology and behaviour as agent technologies evolve.
Our findings have implications for the development of interactive intelligent agents. First, we found that behavioural trust, social alignment, perceived self-agency, and objective task performance were comparable in human-agent and human-human interactions. That is, when assuming human-equivalent roles with matched performance, agents elicit behavioural trust, facilitate effective interaction, and preserve user agency and performance. This indicates that well-performing agents are afforded instrumental value on par with humans, making them promising collaborators in goal-directed tasks. To move beyond instrumental parity towards successful collaboration, our review also reveals key design considerations. Specifically, we found that agents were attributed less responsibility than humans, whereas collaboration typically assumes shared responsibility137, raising concerns that individuals paired with agents may be overburdened with accountability. This necessitates designing robust accountability structures and clear human-agent communication. Another design consideration arises from our review’s focus on agents with human-level performance, which may promote reciprocity and safeguard user agency. However, there are scenarios where leveraging agents’ superhuman capabilities and granting them primary control may be advantageous. Recent AI advances have enabled these agents to surpass humans in both knowledge-based tasks2,4,5 and situational judgment tasks138. Thus, agent design should strategically calibrate agent capabilities to task demands while ensuring transparency and explainability.
Second, we found reduced prosocial behaviour and moral engagement with agent vs. human partners. Agents were also perceived as less competent, likeable, socially present, and attributed less agency and responsibility. From a moral psychology perspective, our findings reflect reduced moral standing for agents, with weaker attributions of both moral agency (capacity to be a source of moral action) and moral patiency (capacity to be a subject of moral concern and obligations)139,140. These deficits suggest that agents are not afforded intrinsic value140 to the same extent as humans. While prior research found that participants compensated an ostracised agent during a Cyberball game141, mirroring behaviour in human-human ostracism, the compensation effect for agents was smaller than that observed for human targets142,143. Accordingly, agents receive some intrinsic value, but to a meaningfully lesser extent. Echoing this interpretation, a meta-analysis of agent anthropomorphism found no significant effect on empathy towards agents74, suggesting that apparent human-likeness alone is insufficient to increase empathic concern. These patterns also reveal a distinction between scholarly arguments that artificial agents could warrant moral consideration133 and empirical evidence that lay individuals interacting with current agents do not afford them equivalent moral standing and intrinsic value. This warrants caution in morally sensitive domains like healthcare, education, and social services, where decisions directly affect welfare and vulnerability144,145, and insufficient prosocial and moral consideration may be consequential. For instance, in AI-mediated tutoring, students may feel less moral obligation to engage honestly or respectfully with the system, potentially increasing dishonest behaviours or disregarding learning guidance. Developers should ensure system benefits outweigh risks of diminished user prosociality and morality and retain human oversight for critical decisions. Design efforts should be put in promoting prosociality/morality in human-agent interaction, potentially leveraging psychological levers like psychological flexibility146 and gratitude147. Operational practices, including robust accountability frameworks, incident logging148, and harm mitigation protocols, should be embedded throughout development cycles.
Whereas agents were generally perceived as possessing fewer social attributes, high heterogeneity in these partner effects highlights important design opportunities. Our review identified moderating effects of several agent and interaction characteristics, indicating that these agent disadvantages are malleable through strategic design and deployment. For example, partner effects on perceived likeability and competence disappeared when appearances were matched (vs. differed). This suggests that calibrating agent appearance relative to human counterparts can improve users’ impressions, which may be particularly relevant when agents serve in customer-facing roles and act as cues that shape corporate brand perceptions149. Additionally, the partner effect on agency attribution disappeared in computer-mediated (vs. face-to-face) interactions, suggesting that the interaction medium can be leveraged to shape users’ mental models of agents in line with design intent—whether as intentional actors (using computer-mediated interfaces) or as tools with limited agency (face-to-face to anchor expectations and prevent over-reliance). The partner effect on responsibility attribution disappeared in real-time (vs. hypothetical) interactions. Agent actions, decisions, and role boundaries should therefore be clearly conveyed through interaction to guide appropriate responsibility attribution. High heterogeneity was also observed in partner effects across interaction experiences, though moderators were identified for only two response types. Regarding perceived partner relational qualities, the partner effect was more pronounced in hypothetical (vs. real-time) interactions and when focused on customer orientation (vs. interactivity/rapport/empathy). This highlights the importance of designing agents to convey relational qualities through direct interaction rather than relying on users’ abstract expectations. Moreover, designers should explicitly signal user-oriented intent through agent behaviour during customer interactions, such as by prioritising user goals and demonstrating alignment with user interests. Finally, interaction satisfaction was higher with agents than with humans in oppositional (vs. cooperative) interactions, suggesting opportunities to design agents as assessors, negotiators, or challengers in scenarios where human opponents may trigger discomfort, evaluation anxiety, or interpersonal friction.
Conclusion
This systematic review and meta-analysis provides a comprehensive understanding of similarities and differences in human responses between human-agent and human-human interactions. Across 23 types of psychological and behavioural responses, we found that individuals exhibited less prosocial behaviour and moral engagement when interacting with agents vs. humans. They also attributed less agency and responsibility to agents, perceiving them as less competent, likeable, and socially present. In contrast, individuals’ social alignment, trust in partners, personal agency, task performance, and interaction experiences were generally comparable when interacting with agents vs. humans. These findings indicate that agents are afforded instrumental value on par with humans yet lack comparable intrinsic value. Furthermore, we identified several moderators across study, participant, partner, interaction, and response characteristics that shaped these partner effects. Overall, this review offers important theoretical insights into human-agent interaction and practical implications for the development of interactive intelligent agents.
Supplementary information
Supplementary information for “A systematic review and meta-analysis of psychological and behavioural responses in human-agent vs. human-human interactions”.
Acknowledgements
We are grateful to the Editor and two Reviewers for their constructive comments and suggestions. We sincerely thank Dr Miguel Vadillo for his guidance on meta-analytic methodology and the development of the research quality checklist, and Dr Brady Roberts and Dr David B. Wilson for their guidance on effect size calculations. We also thank Ms Katharine Thompson and Miss Nicole Urquhart, librarians at Imperial College London, for their support with EndNote and the literature search during the early stages of the review. Finally, we extend our sincere thanks to all authors who responded to our inquiries and shared data and statistical information with us despite their busy schedules. The authors received no specific funding for this work.
Author contributions
J.Z.: Conceptualization, Methodology, Software, Formal analysis, Investigation, Resources, Data curation, Writing – original draft, Writing – review & editing, Visualization, Project administration. F.C.: Methodology, Investigation, Writing – review & editing. J.B.: Methodology, Investigation, Writing – review & editing. T.P.: Conceptualization, Methodology, Investigation, Writing – review & editing, Supervision, Project administration. N.v.Z.: Conceptualization, Methodology, Investigation, Writing – review & editing, Supervision, Project administration.
Peer review
Peer review information
Communications Psychology thanks the anonymous reviewers for their contribution to the peer review of this work. Primary Handling Editors: Troby Ka-Yan Lui. A peer review file is available.
Data availability
The dataset supporting the findings of this study, including calculated effect sizes, moderator values, research quality ratings for individual studies, and associated study materials, is available in the Open Science Framework: 10.17605/OSF.IO/4X26R150.
Code availability
All R code used to calculate effect sizes for individual studies and to perform the frequentist and Bayesian meta-analyses is available via the Open Science Framework: 10.17605/OSF.IO/4X26R151.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
The online version contains supplementary material available at 10.1038/s44271-026-00466-z.
References
- 1.Raees, M., Meijerink, I., Lykourentzou, I., Khan, V.-J. & Papangelis, K. From explainable to interactive AI: A literature review on current trends in human-AI interaction. Int. J. Hum. -Comput. Stud.189, 103301 (2024). [Google Scholar]
- 2.Luo, X. et al. Large language models surpass human experts in predicting neuroscience results. Nat. Hum. Behav.9, 305–315 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.McKinney, S. M. et al. International evaluation of an AI system for breast cancer screening. Nature577, 89–94 (2020). [DOI] [PubMed] [Google Scholar]
- 4.Abdelsamie, M. & Wang, H. Comparative Analysis of LLM-based Market Prediction and Human Expertise with Sentiment Analysis and Machine Learning Integration. In 2024 7th International Conference on Data Science and Information Technology (DSIT) 1–6 10.1109/dsit61374.2024.10881868 (IEEE, Nanjing, China, 2024).
- 5.Bojić, L., Kovačević, P. & Čabarkapa, M. Does GPT-4 surpass human performance in linguistic pragmatics? Humanit. Soc. Sci. Commun. 12, 1–10 (2025).
- 6.Shah, R. et al. An Approach to Technical AGI Safety and Security. Preprint at 10.48550/ARXIV.2504.01849 (2025).
- 7.Ahmed, N. et al. Deep learning-based natural language processing in human–agent interaction: Applications, advancements and challenges. Nat. Lang. Process. J.9, 100112 (2024). [Google Scholar]
- 8.Russell, S. & Norvig, P. Artificial Intelligence: A Modern Approach. (Pearson, Boston, 2016).
- 9.Swarup, S. Agency in the Age of AI. Preprint at 10.48550/ARXIV.2502.00648 (2025).
- 10.Casper, S. et al. The AI Agent Index. Preprint at 10.48550/ARXIV.2502.01635 (2025).
- 11.Zhao, C. & Xu, W. Human-AI Interaction Design Standards. Preprint at 10.48550/ARXIV.2503.16472 (2025).
- 12.Li, J., Yang, Y., Liao, Q. V., Zhang, J. & Lee, Y.-C. As Confidence Aligns: Understanding the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems 1–16 10.1145/3706598.3713336 (ACM, Yokohama Japan, 2025).
- 13.Seeber, I. et al. Machines as teammates: A research agenda on AI in team collaboration. Inf. Manag.57, 103174 (2020). [Google Scholar]
- 14.Wiederhold, B. K. Humanity’s Evolving Conversations: AI as Confidant, Coach, and Companion. Cyberpsychology Behav. Soc. Netw.27, 750–752 (2024). [DOI] [PubMed] [Google Scholar]
- 15.De Vries, K. You never fake alone. Creative AI in action. Inf. Commun. Soc.23, 2110–2127 (2020). [Google Scholar]
- 16.Mao, S. et al. ALYMPICS: LLM agents meet game theory. In Proceedings of the 31st International Conference on Computational Linguistics 2845–2866 (Association for Computational Linguistics, 2025).
- 17.Phaijit, O., Sammut, C. & Johal, W. Let’s Compete! The Influence of Human-Agent Competition and Collaboration on Agent Learning and Human Perception. In Proceedings of the 10th International Conference on Human-Agent Interaction 86–94 10.1145/3527188.3561922 (ACM, Christchurch New Zealand, 2022).
- 18.Mu, C. et al. Multi-agent, human–agent and beyond: A survey on cooperation in social dilemmas. Neurocomputing610, 128514 (2024). [Google Scholar]
- 19.Jiang, T., Sun, Z., Fu, S. & Lv, Y. Human-AI interaction research agenda: A user-centered perspective. Data Inf. Manag.8, 100078 (2024). [Google Scholar]
- 20.Kirk, H. R., Gabriel, I., Summerfield, C., Vidgen, B. & Hale, S. A. Why human–AI relationships need socioaffective alignment. Humanit. Soc. Sci. Commun.12, 728 (2025). [Google Scholar]
- 21.Ben-Zion, Z. Why we need mandatory safeguards for emotionally responsive AI. Nature643, 9–9 (2025). [DOI] [PubMed] [Google Scholar]
- 22.Perez-Osorio, J. & Wykowska, A. Adopting the intentional stance toward natural and artificial agents. Philos. Psychol.33, 369–395 (2020). [Google Scholar]
- 23.Krpan, D., Booth, J. E. & Damien, A. The positive–negative–competence (PNC) model of psychological responses to representations of robots. Nat. Hum. Behav.7, 1933–1954 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Bütepage, J. & Kragic, D. Human-Robot Collaboration: From Psychology to Social Robotics. Preprint at 10.48550/ARXIV.1705.10146 (2017).
- 25.Krämer, N. C., Von Der Pütten, A. & Eimler, S. Human-Agent and Human-Robot Interaction Theory: Similarities to and Differences from Human-Human Interaction. In Human-Computer Interaction: The Agency Perspective (eds Zacarias, M. & De Oliveira, J. V.) vol. 396 215–240 (Springer Berlin Heidelberg, Berlin, Heidelberg, 2012).
- 26.Kempt, H. Artificial Social Agents. in Chatbots and the Domestication of AI 77–135 10.1007/978-3-030-56290-8_5 (Springer International Publishing, Cham, 2020).
- 27.Fox, J. & Gambino, A. Relationship Development with Humanoid Social Robots: Applying Interpersonal Theories to Human–Robot Interaction. Cyberpsychology Behav. Soc. Netw.24, 294–299 (2021). [DOI] [PubMed] [Google Scholar]
- 28.Chugunova, M. & Sele, D. We and It: An interdisciplinary review of the experimental evidence on how humans interact with machines. J. Behav. Exp. Econ.99, 101897 (2022). [Google Scholar]
- 29.Zhang, G., Chong, L., Kotovsky, K. & Cagan, J. Trust in an AI versus a Human teammate: The effects of teammate identity and performance on Human-AI cooperation. Comput. Hum. Behav.139, 107536 (2023). [Google Scholar]
- 30.Alarcon, G. M., Capiola, A., Hamdan, I. A., Lee, M. A. & Jessup, S. A. Differential biases in human-human versus human-robot interactions. Appl. Ergon.106, 103858 (2023). [DOI] [PubMed] [Google Scholar]
- 31.De Visser, E. J. et al. Almost human: Anthropomorphism increases trust resilience in cognitive agents. J. Exp. Psychol. Appl.22, 331–349 (2016). [DOI] [PubMed] [Google Scholar]
- 32.Maehigashi, A., Tsumura, T. & Yamada, S. Experimental Investigation of Trust in Anthropomorphic Agents as Task Partners. In Proceedings of the 10th International Conference on Human-Agent Interaction 302–305 10.1145/3527188.3563921 (ACM, Christchurch New Zealand, 2022).
- 33.Campagna, G. & Rehm, M. A Systematic Review of Trust Assessments in Human–Robot Interaction. ACM Trans. Hum. -Robot Interact.14, 1–35 (2025). [Google Scholar]
- 34.Lee, W.-Y. et al. Interactive Vignettes: Enabling Large-Scale Interactive HRI Research. In 2021 30th IEEE International Conference on Robot & Human Interactive Communication (RO-MAN) 1289–1296 10.1109/RO-MAN50785.2021.9515376 (IEEE, Vancouver, BC, Canada, 2021).
- 35.Greussing, E. et al. Researching interactions between humans and machines: methodological challenges. Publizistik67, 531–554 (2022). [Google Scholar]
- 36.Schecter, A. et al. Vero: An accessible method for studying human–AI teamwork. Comput. Hum. Behav.141, 107606 (2023). [Google Scholar]
- 37.Edwards, A. & Edwards, C. Does the Correspondence Bias Apply to Social Robots?: Dispositional and Situational Attributions of Human Versus Robot Behavior. Front. Robot. AI8, 788242 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Finkel, M. & Krämer, N. C. Humanoid Robots – Artificial. Human-like. Credible? Empirical Comparisons of Source Credibility Attributions Between Humans, Humanoid Robots, and Non-human-like Devices. Int. J. Soc. Robot.14, 1397–1411 (2022). [Google Scholar]
- 39.Basili, P. et al. Inferring the goal of an approaching agent: A human-robot study. In 2012 IEEE RO-MAN: The 21st IEEE International Symposium on Robot and Human Interactive Communication 527–532 10.1109/ROMAN.2012.6343805 (IEEE, Paris, France, 2012).
- 40.Haring, K. et al. I’m Not Playing Anymore! A Study Comparing Perceptions of Robot and Human Cheating Behavior. In Social Robotics (eds Salichs, M. A. et al. vol. 11876 410–419 (Springer International Publishing, Cham, 2019).
- 41.Joosse, M., Lohse, M., Van Berkel, N., Sardar, A. & Evers, V. Making Appearances: How Robots Should Approach People. ACM Trans. Hum. -Robot Interact.10, 1–24 (2021). [Google Scholar]
- 42.Gunes, H. et al. Reproducibility in Human-Robot Interaction: Furthering the Science of HRI. Curr. Robot. Rep.3, 281–292 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Hornbæk, K., Sander, S. S., Bargas-Avila, J. A. & Grue Simonsen, J. Is once enough?: on the extent and content of replications in human-computer interaction. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems 3523–3532 10.1145/2556288.2557004 (ACM, Toronto Ontario Canada, 2014).
- 44.Huang, G. & Wang, S. Is artificial intelligence more persuasive than humans? A meta-analysis. J. Commun.73, 552–562 (2023). [Google Scholar]
- 45.Fox, J. et al. Avatars Versus Agents: A Meta-Analysis Quantifying the Effect of Agency on Social Influence. Hum.–Computer Interact.30, 401–432 (2015). [Google Scholar]
- 46.Fahim, M. A. A., Khan, M. M. H., Jensen, T. & Albayram, Y. Human vs. Automation: Which One Will You Trust More If You Are About to Lose Money? Int. J. Hum.–Computer Interact.39, 2420–2435 (2023). [Google Scholar]
- 47.Wang, C., Li, Y., Fu, W. & Jin, J. Whether to trust chatbots: Applying the event-related approach to understand consumers’ emotional experiences in interactions with chatbots in e-commerce. J. Retail. Consum. Serv.73, 103325 (2023). [Google Scholar]
- 48.Seitz, L., Bekmeier-Feuerhahn, S. & Gohil, K. Can we trust a chatbot like a physician? A qualitative study on understanding the emergence of trust toward diagnostic chatbots. Int. J. Hum. -Comput. Stud.165, 102848 (2022). [Google Scholar]
- 49.Reeves, B. & Nass, C. I. The Media Equation: How People Treat Computers, Television, and New Media like Real People and Places. (CSLI Publications; Cambridge University Press, Stanford, Calif.: New York, 1996).
- 50.Nass, C., Steuer, J. & Tauber, E. R. Computers are social actors. In Proceedings of the SIGCHI conference on Human factors in computing systems 72–78 10.1145/191666.191703 (ACM, Boston, Massachusetts USA, 1994).
- 51.Bailenson, J. N., Blascovich, J., Beall, A. C. & Loomis, J. M. Interpersonal Distance in Immersive Virtual Environments. Pers. Soc. Psychol. Bull.29, 819–833 (2003). [DOI] [PubMed] [Google Scholar]
- 52.Damholdt, M. F., Quick, O. S., Seibt, J., Vestergaard, C. & Hansen, M. A Scoping Review of HRI Research on ‘Anthropomorphism’: Contributions to the Method Debate in HRI. Int. J. Soc. Robot.15, 1203–1226 (2023). [Google Scholar]
- 53.Epley, N., Waytz, A. & Cacioppo, J. T. On seeing human: A three-factor theory of anthropomorphism. Psychol. Rev.114, 864–886 (2007). [DOI] [PubMed] [Google Scholar]
- 54.Fischer, K. Tracking Anthropomorphizing Behavior in Human-Robot Interaction. ACM Trans. Hum. -Robot Interact.11, 1–28 (2022). [Google Scholar]
- 55.Sundar, S. S. The MAIN Model: A Heuristic Approach to Understanding Technology Effects on Credibility. In Digital Media, Youth, and Credibility (eds Metzger, M. J. & Flanagin, A. J.) 73–100 (The MIT Press, Cambridge, MA, 2008).
- 56.Lee, K. M. Presence, Explicated. Commun. Theory14, 27–50 (2004). [Google Scholar]
- 57.Xu, K., Chen, M. & You, L. The Hitchhiker’s Guide to a Credible and Socially Present Robot: Two Meta-Analyses of the Power of Social Cues in Human–Robot Interaction. Int. J. Soc. Robot.15, 269–295 (2023). [Google Scholar]
- 58.Mori, M., MacDorman, K. & Kageki, N. The Uncanny Valley [From the Field]. IEEE Robot. Autom. Mag.19, 98–100 (2012). [Google Scholar]
- 59.Madhavan, P. & Wiegmann, D. A. Similarities and differences between human–human and human–automation trust: an integrative review. Theor. Issues Ergon. Sci.8, 277–301 (2007). [Google Scholar]
- 60.Griffin, E. Chapter 29: The Media Equation of Byron Reeves & Clifford Nass. in A First Look at Communication Theory (McGraw-Hill).
- 61.Lee, E.-J. Minding the source: toward an integrative theory of human–machine communication. Hum. Commun. Res.50, 184–193 (2024). [Google Scholar]
- 62.Lermann Henestrosa, A., Greving, H. & Kimmerle, J. Automated journalism: The effects of AI authorship and evaluative information on the perception of a science journalism article. Comput. Hum. Behav.138, 107445 (2023). [Google Scholar]
- 63.Edwards, C., Edwards, A., Spence, P. R. & Shelton, A. K. Is that a bot running the social media feed? Testing the differences in perceptions of communication quality for a human agent and a bot agent on Twitter. Comput. Hum. Behav.33, 372–376 (2014). [Google Scholar]
- 64.Skjuve, M., Haugstveit, I. M., Følstad, A. & Brandtzaeg, P. B. Help! Is my chatbot falling into the uncanny valley? An empirical study of user experience in human-chatbot interaction. Hum. Technol. 30–54 10.17011/ht/urn.201902201607 (2019).
- 65.Hong, J.-W. & Curran, N. M. Artificial Intelligence, Artists, and Art: Attitudes Toward Artwork Produced by Humans vs. Artificial Intelligence. ACM Trans. Multimed. Comput. Commun. Appl.15, 1–16 (2019). [Google Scholar]
- 66.Schleidgen, S., Friedrich, O., Gerlek, S., Assadi, G. & Seifert, J. The concept of “interaction” in debates on human–machine interaction. Humanit. Soc. Sci. Commun.10, 551 (2023). [Google Scholar]
- 67.Zheng, Q., Tang, Y., Liu, Y., Liu, W. & Huang, Y. UX Research on Conversational Human-AI Interaction: A Literature Review of the ACM Digital Library. In CHI Conference on Human Factors in Computing Systems 1–24 10.1145/3491102.3501855 (ACM, New Orleans LA USA, 2022).
- 68.Motger, Q., Franch, X. & Marco, J. Software-Based Dialogue Systems: Survey, Taxonomy, and Challenges. ACM Comput. Surv.55, 1–42 (2023). [Google Scholar]
- 69.Knote, R., Janson, A., Söllner, M. & Leimeister, J. M. Classifying Smart Personal Assistants: An Empirical Cluster Analysis. in 10.24251/HICSS.2019.245 (2019).
- 70.Fitrianie, S., Bruijnes, M., Richards, D., Bönsch, A. & Brinkman, W.-P. The 19 Unifying Questionnaire Constructs of Artificial Social Agents: An IVA Community Analysis. in Proceedings of the 20th ACM International Conference on Intelligent Virtual Agents 1–8 10.1145/3383652.3423873 (ACM, Virtual Event Scotland UK, 2020).
- 71.Janlert, L.-E. & Stolterman, E. The Meaning of Interactivity—Some Proposals for Definitions and Measures. Hum.–Computer Interact.32, 103–138 (2017). [Google Scholar]
- 72.Blut, M., Wang, C., Wünderlich, N. V. & Brock, C. Understanding anthropomorphism in service provision: a meta-analysis of physical robots, chatbots, and other AI. J. Acad. Mark. Sci.49, 632–658 (2021). [Google Scholar]
- 73.Klein, S. H. The effects of human-like social cues on social responses towards text-based conversational agents—a meta-analysis. Humanit. Soc. Sci. Commun.12, 1322 (2025). [Google Scholar]
- 74.Roesler, E., Manzey, D. & Onnasch, L. A meta-analysis on the effectiveness of anthropomorphism in human-robot interaction. Sci. Robot.6, eabj5425 (2021). [DOI] [PubMed] [Google Scholar]
- 75.Lombard, M. & Xu, K. Social Responses to Media Technologies in the 21st Century: The Media are Social Actors Paradigm. Hum. -Mach. Commun.2, 29–55 (2021). [Google Scholar]
- 76.Are the articles in the IEEE Xplore digital library indexed in SCOPUS. IEEE Online Support Centerhttps://ieeemsd.my.site.com/onlinesupport/s/article/Are-the-articles-in-the-IEEE-Xplore-digital-library-indexed-in-SCOPUS#:~:text=Scopus%20is%20an%20online%20abstract,abstract)%20can%20be%20made%20visible (2021).
- 77.Edwards, C., Edwards, A., Albrehi, F. & Spence, P. Interpersonal impressions of a social robot versus human in the context of performance evaluations. Commun. Educ.70, 165–182 (2021). [Google Scholar]
- 78.Toh, G. et al. Digital interventions for subjective and objective social isolation among individuals with mental health conditions: a scoping review. BMC Psychiatry22, 331 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Nickerson, R. C., Varshney, U. & Muntermann, J. A method for taxonomy development and its application in information systems. Eur. J. Inf. Syst.22, 336–359 (2013). [Google Scholar]
- 80.Johnson, B. T. Toward a more transparent, rigorous, and generative psychology. Psychol. Bull.147, 1–15 (2021). [DOI] [PubMed] [Google Scholar]
- 81.Sterne, J. A. C. et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ l4898 10.1136/bmj.l4898 (2019). [DOI] [PubMed]
- 82.Whiting, P. et al. ROBIS: A new tool to assess risk of bias in systematic reviews was developed. J. Clin. Epidemiol.69, 225–234 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Hong, Q. N. et al. The Mixed Methods Appraisal Tool (MMAT) version 2018 for information professionals and researchers. Educ. Inf.34, 285–291 (2018). [Google Scholar]
- 84.Aromataris, E. & Munn, Z. JBI Manual for Evidence Synthesis. (Joanna Briggs Institute, Adelaide, Australia, 2020).
- 85.Harrison, R., Jones, B., Gardner, P. & Lawton, R. Quality assessment with diverse studies (QuADS): an appraisal tool for methodological and reporting quality in systematic reviews of mixed- or multi-method studies. BMC Health Serv. Res21, 144 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Kmet, L. M., Lee, R. C. & Cook, L. S. Standard Quality Assessment Criteria for Evaluating Primary Research Papers from a Variety of Fields. (Alberta Heritage Foundation for Medical Research, Edmonton, Alta., 2004).
- 87.Rodríguez-Prada, C., Burgaleta, M., Morís Fernández, L., Vadillo, M. A. & Soto-Faraco, S. Online Interventions for Mental Health in Times of COVID: A Systematic Review and Quality Assessment of Scientific Production. Collabra Psychol.9, 90197 (2023). [Google Scholar]
- 88.Ferrero, M., Vadillo, M. A. & León, S. P. A valid evaluation of the theory of multiple intelligences is not yet possible: Problems of methodological quality for intervention studies. Intelligence88, 101566 (2021). [Google Scholar]
- 89.Cochrane Handbook for Systematic Reviews of Interventions. (Wiley-Blackwell, Hoboken, NJ, 2019).
- 90.Jüni, P. The Hazards of Scoring the Quality of Clinical Trials for Meta-analysis. JAMA282, 1054 (1999). [DOI] [PubMed] [Google Scholar]
- 91.Liu, R. T. et al. A systematic review and Bayesian meta-analysis of 30 years of stress generation research: Clinical, psychological, and sociodemographic risk and protective factors for prospective negative life events. Psychol. Bull.150, 1021–1069 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Jackson, D. & Turner, R. Power analysis for random-effects meta-analysis. Res. Synth. Methods8, 290–302 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Borenstein, M., Hedges, L. V., Higgins, J. P. T. & Rothstein, H. R. A basic introduction to fixed-effect and random-effects models for meta-analysis. Res. Synth. Methods1, 97–111 (2010). [DOI] [PubMed] [Google Scholar]
- 94.Assink, M. & Wibbelink, C. J. M. Fitting three-level meta-analytic models in R: A step-by-step tutorial. Quant. Methods Psychol.12, 154–174 (2016). [Google Scholar]
- 95.Viechtbauer, W. Conducting Meta-Analyses in R with the metafor Package. J. Stat. Softw. 36, 1–48 (2010).
- 96.Borenstein, M., Hedges, L. V., Higgins, J. P. T. & Rothstein, H. R. Introduction to Meta-Analysis. 10.1002/9780470743386 (Wiley, 2009).
- 97.Yang, Y. & Konrath, S. A systematic review and meta-analysis of the relationship between economic inequality and prosocial behaviour. Nat. Hum. Behav.7, 1899–1916 (2023). [DOI] [PubMed] [Google Scholar]
- 98.Nakagawa, S. et al. orchaRd 2.0: An R package for visualising meta-analyses with orchard plots. Methods Ecol. Evol.14, 2003–2010 (2023). [Google Scholar]
- 99.Pustejovsky, J. E. & Tipton, E. Small-Sample Methods for Cluster-Robust Variance Estimation and Hypothesis Testing in Fixed Effects Models. J. Bus. Econ. Stat.36, 672–683 (2018). [Google Scholar]
- 100.Viechtbauer, W. & Cheung, M. W.-L. Outlier and influence diagnostics for meta-analysis. Res. Synth. Methods1, 112–125 (2010). [DOI] [PubMed] [Google Scholar]
- 101.Bürkner, P.-C. brms: An R Package for Bayesian Multilevel Models Using Stan. J. Stat. Softw. 80, 1–28 (2017).
- 102.Gronau, Q. F., Singmann, H. & Wagenmakers, E.-J. bridgesampling: An R Package for Estimating Normalizing Constants. J. Stat. Softw. 92, 1–29 (2020).
- 103.Thompson, C. G. & Semma, B. An alternative approach to frequentist meta-analysis: A demonstration of Bayesian meta-analysis in adolescent development research. J. Adolesc.82, 86–102 (2020). [DOI] [PubMed] [Google Scholar]
- 104.Gronau, Q. F. et al. A Bayesian model-averaged meta-analysis of the power pose effect with informed and default priors: the case of felt power. Compr. Results Soc. Psychol.2, 123–138 (2017). [Google Scholar]
- 105.Van Erp, S., Verhagen, J., Grasman, R. P. P. P. & Wagenmakers, E.-J. Estimates of Between-Study Heterogeneity for 705 Meta-Analyses Reported in Psychological Bulletin From 1990–2013. J. Open Psychol. Data5, 4 (2017).
- 106.Williams, D. R., Rast, P. & Bürkner, P.-C. Bayesian Meta-Analysis with Weakly Informative Prior Distributions. Preprint at 10.31234/osf.io/7tbrm (2018).
- 107.Gelman, A. & Rubin, D. B. Inference from Iterative Simulation Using Multiple Sequences. Stat. Sci. 7, 457–472 (1992).
- 108.Wetzels, R. & Wagenmakers, E.-J. A default Bayesian hypothesis test for correlations and partial correlations. Psychon. Bull. Rev.19, 1057–1064 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Maassen, E., Van Assen, M. A. L. M., Nuijten, M. B., Olsson-Collentine, A. & Wicherts, J. M. Reproducibility of individual effect sizes in meta-analyses in psychology. PLOS ONE15, e0233107 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Viechtbauer, W. Assembling Data for a Meta-Analysis of Standardized Mean Differences. The metafor Package: A Meta-Analysis Package for Rhttps://www.metafor-project.org/doku.php/tips:assembling_data_smd#fnt__1.
- 111.Pfattheicher, S., Nielsen, Y. A. & Thielmann, I. Prosocial behavior and altruism: A review of concepts and definitions. Curr. Opin. Psychol.44, 124–129 (2022). [DOI] [PubMed] [Google Scholar]
- 112.Eisenberg, N. Emotion, Regulation, and Moral Development. Annu. Rev. Psychol.51, 665–697 (2000). [DOI] [PubMed] [Google Scholar]
- 113.Aquino, K., Freeman, D., Reed, A., Lim, V. K. G. & Felps, W. Testing a social-cognitive model of moral behavior: The interactive influence of situations and moral identity centrality. J. Pers. Soc. Psychol.97, 123–141 (2009). [DOI] [PubMed] [Google Scholar]
- 114.Fiske, S. T. Stereotype Content: Warmth and Competence Endure. Curr. Dir. Psychol. Sci.27, 67–73 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Hubner, A. Y. & Bullock, O. M. Why Science Should Have a Female Face: Female Experts Increase Liking, Competence, and Trust in Science. Sci. Commun. 10755470241295676 10.1177/10755470241295676 (2024).
- 116.Sanders, T. The Likeability Factor: How to Boost Your L-Factor & Achieve Your Life’s Dreams. (Crown, New York, 2006).
- 117.Zheng, T., Duan, X., Zhang, K., Yang, X. & Jiang, Y. How Chatbots’ Anthropomorphism Affects User Satisfaction: The Mediating Role of Perceived Warmth and Competence. In E-Business. Digital Empowerment for an Intelligent Future (eds Tu, Y. & Chi, M.) vol. 481 96–107 (Springer Nature Switzerland, Cham, 2023).
- 118.Moruzzi, C. Creative Agents: Rethinking Agency and Creativity in Human and Artificial Systems. J. Aesthet. Phenomenol.9, 245–268 (2022). [Google Scholar]
- 119.Li, X., Hatada, Y. & Narumi, T. It’s My Fingers’ Fault: Investigating the Effect of Shared Avatar Control on Agency and Responsibility Attribution. IEEE Trans. Vis. Comput. Graph.31, 2859–2869 (2025). [DOI] [PubMed] [Google Scholar]
- 120.Mayer, R. C., Davis, J. H. & Schoorman, F. D. An Integrative Model of Organizational Trust. Acad. Manag. Rev.20, 709 (1995). [Google Scholar]
- 121.Kopp, S. Social resonance and embodied coordination in face-to-face conversation with artificial interlocutors. Speech Commun.52, 587–597 (2010). [Google Scholar]
- 122.Morin, A. Toward a Glossary of Self-related Terms. Front. Psychol. 8, 280 (2017). [DOI] [PMC free article] [PubMed]
- 123.Zoogah, D. B. Theoretical Perspectives of Strategic Behavior. in Strategic Followership 17–37 10.1057/9781137354426_2 (Palgrave Macmillan US, New York, 2014).
- 124.Fatima, J. K. & Razzaque, M. A. Service quality and satisfaction in the banking sector. Int. J. Qual. Reliab. Manag.31, 367–379 (2014). [Google Scholar]
- 125.Hamilton, M., Kaltcheva, V. D. & Rohm, A. J. Social Media and Value Creation: The Role of Interaction Satisfaction and Interaction Immersion. J. Interact. Mark.36, 121–133 (2016). [Google Scholar]
- 126.Rodgers, M. A. & Pustejovsky, J. E. Evaluating meta-analytic methods to detect selective reporting in the presence of dependent effect sizes. Psychol. Methods26, 141–160 (2021). [DOI] [PubMed] [Google Scholar]
- 127.Fitrianie, S., Bruijnes, M., Richards, D., Abdulrahman, A. & Brinkman, W.-P. What are We Measuring Anyway?: - A Literature Survey of Questionnaires Used in Studies Reported in the Intelligent Virtual Agent Conferences. In Proceedings of the 19th ACM International Conference on Intelligent Virtual Agents 159–161 10.1145/3308532.3329421 (ACM, Paris France, 2019).
- 128.March, C. The Behavioral Economics of Artificial Intelligence: Lessons from Experiments with Computer Players. (Bamberg Economic Research Group, Bamberg University, Bamberg, 2019).
- 129.Van Lissa, C. J., Van Erp, S. & Clapper, E. Selecting relevant moderators with Bayesian regularized meta-regression. Res. Synth. Methods14, 301–322 (2023). [DOI] [PubMed] [Google Scholar]
- 130.Henrich, J., Heine, S. J. & Norenzayan, A. The weirdest people in the world?. Behav. Brain Sci.33, 61–83 (2010). [DOI] [PubMed] [Google Scholar]
- 131.Li, H., Zhang, J. & Huang, K. Meta-Analyzing the Trust-Performance Link in Collaboration: Moderating Effects of Conceptual and Contextual Factors. Public Perform. Manag. Rev.48, 1–34 (2025). [Google Scholar]
- 132.Esterwood, C., Ye, X., Guan, R. & Robert, L. P. Physically Present or Virtually Represented: A Meta-Analysis of Robot Representation Type’s Impact on Acceptance in Human-Robot Interaction. Preprint at 10.31234/osf.io/4ekbm_v1 (2025).
- 133.Harris, J. & Anthis, J. R. The Moral Consideration of Artificial Entities: A Literature Review. Sci. Eng. Ethics27, 53 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 134.Yin, Y., Jia, N. & Wakslak, C. J. AI can help people feel heard, but an AI label diminishes this impact. Proc. Natl. Acad. Sci.121, e2319112121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 135.Drouin, M., Sprecher, S., Nicola, R. & Perkins, T. Is chatting with a sophisticated chatbot as good as chatting online or FTF with a stranger? Comput. Hum. Behav.128, 107100 (2022). [Google Scholar]
- 136.Schneiders, E. et al. Objection Overruled! Lay People can Distinguish Large Language Models from Lawyers, but still Favour Advice from an LLM. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems 1–14 10.1145/3706598.3713470 (ACM, Yokohama Japan, 2025).
- 137.Puerta-Beldarrain, M. et al. A Multifaceted Vision of the Human-AI Collaboration: A Comprehensive Review. IEEE Access13, 29375–29405 (2025). [Google Scholar]
- 138.Mittelstädt, J. M., Maier, J., Goerke, P., Zinn, F. & Hermes, M. Large language models can outperform humans in social situational judgments. Sci. Rep.14, 27449 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 139.Formosa, P., Hipólito, I. & Montefiore, T. AI and the Relationship Between Agency, Autonomy, and Moral Patiency. In Reconfiguring Human Autonomy (eds Anzalone, M., Achella, S., Battaglia, F. & Donise, A.) vol. 40 35–50 (Springer Nature Switzerland, Cham, 2026).
- 140.Floridi, L. On the intrinsic value of information objects and the infosphere. Ethics Inf. Technol.4, 287–304 (2002). [Google Scholar]
- 141.Zhou, J., Porat, T. & Van Zalk, N. Humans Mindlessly Treat AI Virtual Agents as Social Beings, but This Tendency Diminishes Among the Young: Evidence From a Cyberball Experiment. Hum. Behav. Emerg. Technol.2024, 8864909 (2024). [Google Scholar]
- 142.Sellaro, R., Steenbergen, L., Verkuil, B., Van IJzendoorn, M. H. & Colzato, L. S. Transcutaneous Vagus Nerve Stimulation (tVNS) does not increase prosocial behavior in Cyberball. Front. Psychol.06, Article 499 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 143.Van Der Meulen, M., Van IJzendoorn, M. H. & Crone, E. A. Neural Correlates of Prosocial Behavior: Compensating Social Exclusion in a Four-Player Cyberball Game. PLOS ONE11, e0159045 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 144.Fu, Y. & Weng, Z. Navigating the ethical terrain of AI in education: A systematic review on framing responsible human-centered AI practices. Comput. Educ. Artif. Intell.7, 100306 (2024). [Google Scholar]
- 145.Cecez-Kecmanovic, D. Ethics in the world of automated algorithmic decision-making – A Posthumanist perspective. Inf. Organ.35, 100587 (2025). [Google Scholar]
- 146.Gloster, A. T., Rinner, M. T. B. & Meyer, A. H. Increasing prosocial behavior and decreasing selfishness in the lab and everyday life. Sci. Rep. 10, 21220 (2020). [DOI] [PMC free article] [PubMed]
- 147.Ma, L. K., Tunney, R. J. & Ferguson, E. Does gratitude enhance prosociality?: A meta-analytic review. Psychol. Bull.143, 601–635 (2017). [DOI] [PubMed] [Google Scholar]
- 148.McGregor, S. Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database. Proc. AAAI Conf. Artif. Intell.35, 15458–15463 (2021). [Google Scholar]
- 149.Wünderlich, N. V., Blut, M. & Brock, C. Enhancing corporate brands through service robots: The impact of anthropomorphic design metaphors on corporate brand perceptions. J. Prod. Innov. Manag.41, 1022–1046 (2024). [Google Scholar]
- 150.Zhou, J., Corbett, F., Byun, J., Porat, T. & Van Zalk, N. Dataset for ‘Psychological and behavioural responses in human-agent vs. human-human interactions’. OSF 10.17605/OSF.IO/4X26R (2025). [DOI] [PMC free article] [PubMed]
- 151.Zhou, J., Corbett, F., Byun, J., Porat, T. & Van Zalk, N. R Code for ‘Psychological and behavioural responses in human-agent vs. human-human interactions’. 10.17605/OSF.IO/4X26R (2025). [DOI] [PMC free article] [PubMed]
- 152.Phillips, E., Zhao, X., Ullman, D. & Malle, B. F. What is Human-like?: Decomposing Robots’ Human-like Appearance Using the Anthropomorphic roBOT (ABOT) Database. In Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction 105–113 10.1145/3171221.3171268 (ACM, Chicago IL USA, 2018).
- 153.Biocca, F. The Cyborg’s Dilemma: Progressive Embodiment in Virtual Environments. J. Comput.-Mediat. Commun. 3, JCMC324 (1997).
- 154.Schwarz, K. A., Tonn, S., Büttner, J., Kunde, W. & Pfister, R. Sense of agency in social hierarchies. J. Exp. Psychol. Gen.152, 2957–2976 (2023). [DOI] [PubMed] [Google Scholar]
- 155.Johnson, D. G. & Verdicchio, M. AI, agency and responsibility: the VW fraud case and beyond. AI Soc.34, 639–647 (2019). [Google Scholar]
- 156.Ahmed, A. M. & Salas, O. The Relationship between Behavioral and Attitudinal Trust: A Cross-cultural Study. Rev. Soc. Econ.67, 457–482 (2009). [Google Scholar]
- 157.Rousseau, D. M., Sitkin, S. B., Burt, R. S. & Camerer, C. Not So Different After All: A Cross-Discipline View Of Trust. Acad. Manag. Rev.23, 393–404 (1998). [Google Scholar]
- 158.Wen, W. & Imamizu, H. The sense of agency in perception, behaviour and human–machine interactions. Nat. Rev. Psychol.1, 211–222 (2022). [Google Scholar]
- 159.Self-Disclosure. (Sage Publications, Newbury Park, 1993).
- 160.Krueger, J. I., Heck, P. R., Evans, A. M. & DiDonato, T. E. Social game theory: Preferences, perceptions, and choices. Eur. Rev. Soc. Psychol.31, 222–253 (2020). [Google Scholar]
- 161.Bradley, M. M. & Lang, P. J. Affective Norms for English Words (ANEW): Instruction Manual and Affective Ratings. (1999).
- 162.Horstmann, A. C., Gratch, J. & Krämer, N. C. I just wanna blame somebody, not something! Reactions to a computer agent giving negative feedback based on the instructions of a person. Int. J. Hum. -Comput. Stud.154, 102683 (2021). [Google Scholar]
- 163.Memarian, B. & Mitropoulos, P. Production practices affecting worker task demands in concrete operations: A case study. Work53, 535–550 (2016). [DOI] [PubMed] [Google Scholar]
- 164.Horrey, W. J., Lesch, M. F., Garabet, A., Simmons, L. & Maikala, R. Distraction and task engagement: How interesting and boring information impact driving performance and subjective and physiological responses. Appl. Ergon.58, 342–348 (2017). [DOI] [PubMed] [Google Scholar]
- 165.Mozgai, S., et al. vol. 10498 283–286 (Springer International Publishing, Cham, 2017).
- 166.Gratch, J., et al. vol. 10011 283–294 (Springer International Publishing, Cham, 2016).
- 167.Xu, J., Bryant, D. G. & Howard, A. Would You Trust a Robot Therapist? Validating the Equivalency of Trust in Human-Robot Healthcare Scenarios. In 2018 27th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN) 442–447 10.1109/ROMAN.2018.8525782 (IEEE, Nanjing, 2018).
- 168.Xu, J., Bryant, D. G., Chen, Y.-P. & Howard, A. Robot therapist versus human therapist: Evaluating the effect of corrective feedback on human motor performance. In 2018 International Symposium on Medical Robotics (ISMR) 1–6 10.1109/ISMR.2018.8333308 (IEEE, Atlanta, GA, USA, 2018).
- 169.Liu, Y., Yan, W., Hu, B., Lin, Z. & Song, Y. Chatbots or Humans? Effects of Agent Identity and Information Sensitivity on Users’ Privacy Management and Behavioral Intentions: A Comparative Experimental Study between China and the United States. Int. J. Hum.–Computer Interact.40, 5632–5647 (2024). [Google Scholar]
- 170.De Melo, C. M. & Terada, K. Cooperation with autonomous machines through culture and emotion. PLOS ONE14, e0224758 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary information for “A systematic review and meta-analysis of psychological and behavioural responses in human-agent vs. human-human interactions”.
Data Availability Statement
The dataset supporting the findings of this study, including calculated effect sizes, moderator values, research quality ratings for individual studies, and associated study materials, is available in the Open Science Framework: 10.17605/OSF.IO/4X26R150.
All R code used to calculate effect sizes for individual studies and to perform the frequentist and Bayesian meta-analyses is available via the Open Science Framework: 10.17605/OSF.IO/4X26R151.
















