Significance
How will people’s increasing reliance on large language models (LLMs) influence their opinions about important moral and societal decisions? Our experiments demonstrate that the decisions and advice of LLMs are systematically biased against doing anything, and this bias is stronger than in humans. Moreover, we identified a bias in LLMs’ responses that has not been found in people. LLMs tend to answer “no,” thus flipping their decision/advice depending on how the question is worded. We present some evidence that suggests both biases are induced when fine-tuning LLMs for chatbot applications. These findings suggest that the uncritical reliance on LLMs could amplify and proliferate problematic biases in societal decision-making.
Keywords: large language models, moral decision-making, ethical AI, moral dilemmas
Abstract
As large language models (LLMs) become more widely used, people increasingly rely on them to make or advise on moral decisions. Some researchers even propose using LLMs as participants in psychology experiments. It is, therefore, important to understand how well LLMs make moral decisions and how they compare to humans. We investigated these questions by asking a range of LLMs to emulate or advise on people’s decisions in realistic moral dilemmas. In Study 1, we compared LLM responses to those of a representative U.S. sample (N = 285) for 22 dilemmas, including both collective action problems that pitted self-interest against the greater good, and moral dilemmas that pitted utilitarian cost–benefit reasoning against deontological rules. In collective action problems, LLMs were more altruistic than participants. In moral dilemmas, LLMs exhibited stronger omission bias than participants: They usually endorsed inaction over action. In Study 2 (N = 474, preregistered), we replicated this omission bias and documented an additional bias: Unlike humans, most LLMs were biased toward answering “no” in moral dilemmas, thus flipping their decision/advice depending on how the question is worded. In Study 3 (N = 491, preregistered), we replicated these biases in LLMs using everyday moral dilemmas adapted from forum posts on Reddit. In Study 4, we investigated the sources of these biases by comparing models with and without fine-tuning, showing that they likely arise from fine-tuning models for chatbot applications. Our findings suggest that uncritical reliance on LLMs’ moral decisions and advice could amplify human biases and introduce potentially problematic biases.
As chatbots based on large language models (LLMs) become more widely used across many different contexts (1), the extent of their usefulness in various decisions is increasingly questioned (2). One key concern is the quality of their moral decisions and advice. Are AI systems, such as chatbots, able to make sound moral judgments and decisions?
Moral issues inevitably appear in conversations with chatbots due to their prevalence in everyday scenarios people want advice on (e.g., “Should I tell my friend that their cooking tastes bad, even though it would hurt their feelings?”). For this reason, LLM developers include moral specifications in guidelines to shape the models’ behavior (e.g., ref. 3). ChatGPT, for example, is programmed to “encourage fairness and kindness, and discourage hate,” and not promote illegal activity (3). However, LLMs can display unpredictable, erroneous, or unreliable behavior, such as “hallucinations” (4) and cognitive biases (5).
Moral decisions made by LLMs and other AI agents can have important practical consequences. For example, chatbots can be integrated into autonomous vehicles for decision-making (6–9), which raises the question of how AI agents can—or should—make life-or-death decisions between prioritizing the safety of passengers or sacrificing them for the greater good (e.g., to protect a larger number of pedestrians). Further, LLM chatbots’ moral advice can influence people’s decisions in everyday interactions (10, 11). Recent studies have shown that AI decision-making can be distorted by irrelevant details [e.g., in image classification (12, 13)], and there have been inconsistent findings about AI reasoning and decision-making abilities (14–18). In moral contexts, the potential negative consequences caused by erroneous judgments and decisions could be particularly catastrophic.
One way of studying LLMs is to use methods designed to investigate human psychology (19–22). In this paper, we apply this approach to moral reasoning and decision-making by using an experiment designed to investigate cognitive biases in human moral decision-making. We compare LLM responses to those of human participants and explore systematic similarities and differences. We focus on two possible conflicts related to moral and altruistic decision-making: 1) deciding which action to take when different actions are implied by different moral views, typically “utilitarianism” vs. “deontology,” and 2) self–other trade-offs, where people need to allocate limited resources between themselves vs. others who would benefit more (or in-group members vs. out-group members).
“Utilitarian” vs. “deontological” decision-making is often studied using dilemmas with two opposing choices that align with these two competing moral perspectives (for a review, see ref. 23). According to utilitarianism, actions should be evaluated based on their anticipated consequences for everyone’s well-being, and morally good actions maximize happiness. According to deontology, actions should be evaluated solely based on whether they follow moral rules or norms. A well-known moral dilemma that pits these views against each other is the “trolley problem” (24, 25), where one must decide between letting a runaway trolley run over five people tied to the track or pulling a lever to redirect the trolley so it runs over one person instead (24). Here, the “utilitarian” choice is to pull the lever (i.e., sacrificing one to save many), whereas the “deontological” choice is to do nothing (i.e., upholding the principle of doing no harm).
Recent studies have explored moral decision-making by prominent LLMs such as ChatGPT using variations of the trolley problem (7, 10, 26), often demonstrating systematic differences between responses of LLMs and those of human participants (7, 26). For example, while people endorse sacrificial harm (e.g., killing one person) when the anticipated benefits are large enough (e.g., saving countless lives in the future), the studied LLMs did not (26). These differences imply that LLMs may not make moral decisions in the same way that people do (27). Despite this, there is some evidence that advice given by LLMs influences people’s moral decisions in these types of dilemmas (10).
One limitation of this prior work is that it mainly used trolley problems. Given that trolley problems are highly unrealistic and sometimes absurd (28, 29), these findings might not generalize to the real-world moral dilemmas for which people seek advice (30). Another issue is that trolley problems are likely common in the training data. Even if LLMs make sensible decisions in trolley problems, they might make puzzling and consequential mistakes in novel, naturalistic situations. While some studies have begun to address this gap by using new LLM-generated trolley problems (7), these scenarios are only slight variations of the classic dilemma and do not address the lack of realism.
Further, the simple dissociation between utilitarian and deontological choices in these unrealistic scenarios also fails to acknowledge a central feature of moral decision-making: the presence of uncertainty (31, 32). The trolley problem only includes a brief description of the scenario followed by the two choices. It does not acknowledge that the action might fail to achieve its intended consequence, nor the risk of additional unintended consequences. Due to this uncertainty, the “utilitarian” choice in such dilemmas might fail to achieve the greater good in the real world. Instead, following the deontological rule might often bring about better outcomes that maximize utility (33) [similar to rule utilitarianism (34)]. Therefore, the choice to commit sacrificial harm can more accurately be described as a choice resulting from explicit cost–benefit reasoning (CBR), rather than that it would necessarily lead to better outcomes (31). In this article, we thus label the choices as the “CBR option” vs. the “rule option,” as opposed to utilitarian vs. deontological. Importantly, CBR refers to a “naive” cost–benefit reasoning, which simply counts up the number of people affected under the different outcomes and the associated probabilities rather than taking all possible indirect consequences into account (31). The latter is intractable in real-world situations (33). By the “rule option,” we refer to the choice option that is consistent with following (or not violating) a moral rule (31).
Additionally, in typical trolley problems, actions usually correspond to the CBR option, and omissions to the rule option (35). This confound may significantly impact how results are interpreted. When people choose not to push the man off the footbridge, it is not necessarily because they believe deontology trumps utilitarianism. Instead, it could be due to preferring inaction in situations that are controversial or ambiguous (36–38). There is strong evidence that people prefer causing harm by inaction vs. action [i.e., omission bias (39–41)].
Because of these considerations, we use the moral dilemmas that (31) adapted from ref. 42. These relatively novel dilemmas are based on real-life (sometimes historical) events to make them more believable. They also include both scenarios where following the rule is framed as the action under consideration and others where the CBR option is the action under consideration.
In addition to the moral dilemmas reviewed so far, people often face collective action problems (43–46)—decisions that pit narrow self-interest (or the interests of one’s in-group) against the greater good. These are situations where the incentives for each individual member of a group are misaligned with the interests of everyone involved: If everyone chooses what is best for them individually, then everyone will be worse off, but if they forgo some personal benefits to cooperate, then everyone will be better off. These types of problems take many forms in real-life contexts, such as in the management of natural resources (46) [p.11].
A classic example is the tragedy of the commons (47): If everyone exploits a limited shared resource (e.g., water during a drought) to their maximum immediate benefit, then the resource might be used up faster than it can be replenished. As a result, the group may unnecessarily extinguish the resource to its own detriment. In contrast, if everyone cooperatively limits their consumption to preserve the resource, then everyone can benefit from using it indefinitely. Other examples of the realistic collective action problems we use include decisions about how much money to donate to people in greater need, whether to help a competitor improve their performance, and whether to take personal risks to blow the whistle on corporate wrongdoing (48, 49). In the dilemmas that we use, the self-sacrifice is usually relatively small in comparison to the benefit to the welfare of others. To our knowledge, no prior work has assessed the quality of LLMs’ decisions and advice in this type of problem.
An important consideration is what role LLMs should play in moral decision-making. One application would be to use LLMs instead of humans as participants in psychology experiments (50, 51), given some evidence of similarities between them (e.g., refs. 52, 53). Another application would be to advise people on how to navigate moral dilemmas. Consistent with this idea, people rated ChatGPT’s moral justifications and advice more favorably than that of a representative sample of Americans and the New York Times column “The Ethicist,” suggesting that the models are perceived to be experts in moral decision-making (11). However, it is unclear whether the perceived quality of LLM advice is a good measure of performance or expertise in moral decision-making tasks. This is especially the case given that LLMs are trained through RLHF to give answers that the user likes (54) rather than, for example, answers that consistently align with moral principles. A further criticism of LLMs’ performance is that they are severely limited in their ability to reflect psychological variation across a diverse human population (2, 55, 56), which has particularly negative effects on marginalized groups (57). This limits how effective they can be both as participants in psychology experiments and as advisors.
In this article, we empirically investigate how good LLMs are at predicting and advising people’s decisions in realistic moral dilemmas and collective action problems. Across four experiments, we found that in moral dilemmas, LLMs have a general tendency to 1) answer “no” (“yes–no bias”), and 2) endorse inaction over action (omission bias), whereas for collective action problems, LLMs showed increased levels of altruism. Additionally, we compared versions of the Llama 3.1 model with different types of fine-tuning; results suggest that fine-tuning models for chatbot applications can induce the yes–no bias and amplify the omission bias.
Overall, our results demonstrate that LLMs and people systematically differ in their moral decisions, and some of these deviations can be problematic (e.g., the yes–no bias). We discuss the implications of our findings for what role LLMs should play in moral decision-making.
Results
Study 1.
In this study, we compared the moral decisions of GPT-4-turbo, GPT-4o, Llama 3.1-Instruct, and Claude 3.5 Sonnet (hereafter Claude 3.5) to the responses of a representative sample of U.S. participants recruited on Prolific (N = 285; see Materials and Methods for details). For LLMs, we also explored whether an advice-giving prompt vs. a prompt to answer as an experimental participant affects responses. We gave participants and LLMs 13 moral dilemmas and nine collective action problems. Participants viewed all dilemmas in a randomized order in a within-subjects design.* The LLMs were asked about each scenario individually. We ran 500 iterations of each vignette with each of the models (with exceptions; see summary of valid responses in the online materials).
Decisions in moral dilemmas.
We showed participants and LLMs 13 moral dilemmas (31, 42) between the “CBR option,” which is the (naive) CBR endorsed choice of committing sacrificial harm or breaking a moral rule for the greater good, and the “rule option,” which is the choice of following (or not violating) a moral rule.
Comparisons between LLM and participant responses.
Fig. 1 visualizes the responses given by participants and LLMs (with participant prompt). In individual dilemmas (Panel B), the LLMs either almost always or almost never chose the CBR option. For participants, the responses tended to shrink more toward 0.5.
Fig. 1.
Compared to people, LLMs (with participant prompt) are more influenced by action/omission framing in Study 1. Panel (A) shows the average omission bias for humans and all models, and Panel (B) shows responses for each vignette. In both panels, the red bars show responses for vignettes where the CBR option coincides with action (“Action Framing”), and the blue bars show responses for vignettes where the CBR option coincides with omission (“Omission Framing”). Similar results with the advice-giving prompt is in SI Appendix, Fig. S1.
Only GPT-4o showed a significant Pearson correlation between LLM and participant responses (; SI Appendix, Table S1). We found significant strong correlations between each model’s decision and its advice (GPT-4-turbo: r = 0.85; GPT-4o: r = 0.92; Llama 3.1-Instruct: r = 0.998; Claude 3.5 = r = 0.98; all P < 0.001).
Effect of action and omission framing.
In our vignettes, we mitigated the typical confounding of action with the CBR option and omission with the rule option. In some vignettes, following the rule coincided with action: For instance, in “Veterinarian,” the reader must decide whether to quit their job in which they use animal testing to develop a vaccine that would likely save the lives of many more animals. Here, the rule option coincides with action (quit the job to avoid directly harming animals), and the CBR option coincides with omission (continue the job and save more animals).
As shown in Fig. 1A, all LLMs responded differently depending on how the vignettes were framed: They were less likely to choose the CBR option when it coincided with action than when it coincided with omission, 53% vs. 97%, (GPT-4-turbo: 40% vs. 100%, ; GPT-4o: 70% vs. 100%, ; Llama 3.1-Instruct: 59% vs. 90%, ; Claude 3.5: 42% vs. 100%, ). We found similar results with the advice-giving prompt (SI Appendix, Fig. S1).
Participants were also less likely to choose the CBR option when it coincided with action than when it coincided with omission (50% vs. 55%, ). Overall, this bias was much larger in LLMs than in humans both in terms of the percentage change (difference in LLMs: 45% vs. difference in humans: 5%) and the beta-coefficient (difference in LLMs: vs. difference in humans: ). The difference between human and LLM responses can also be seen in the individual dilemmas (Fig. 1B).
Collective action problems.
Fig. 2A compares the average responses given by humans and LLMs for the collective action problems. To account for the difference in slider range between different vignettes, we calculated an altruism score that normalizes the responses for each dilemma by the range of available options. The altruism score is a number between 0 and 1, where 0 denotes the most selfish response possible in that scenario, and 1 the most altruistic.
Fig. 2.
LLMs are more altruistic than participants in Study 1 (with participant and advice-giving prompts). Panel (A) shows mean altruism scores across humans and all models, and Panel (B) shows mean scores for each vignette. Llama 3.1-Instruct did not respond to the “Drought” vignette. Error bars indicate 95% CI.
As shown in Fig. 2, all LLMs’ decisions and advice were more altruistic than people’s decisions (all P < 0.001; SI Appendix, Table S2). We found significant correlations between the GPT models and participants’ responses with the participant prompt (GPT-4-turbo: ; GPT-4o: ; SI Appendix, Table S3), and strong correlations between each model’s decisions and its advice (GPT-4-turbo: r = 0.88; GPT-4o: r = 0.97; Llama 3.1-Instruct: r = 0.99; Claude 3.5: r = 0.89; all P < 0.001).
Study 2.
Study 2 was a preregistered (https://osf.io/t4w9g) follow-up to investigate the strong omission bias in moral dilemmas found in Study 1. In Study 1, answering “yes” to the question posed in the moral dilemma corresponded with taking action, regardless of whether that action coincided with CBR or rules. Therefore, two explanations for our findings are 1) that LLMs have a general tendency to answer “no” (yes–no bias), and 2) endorse omission over action (omission bias). To test this, in Study 2, we reframed the dilemmas and compared responses of 490 participants from a U.S. sample to the same LLMs as in Study 1.†
We used six moral dilemmas from Study 1 that can be reasonably reframed to test for both biases. To illustrate this, consider the “Assisted Suicide” vignette (Table 1). In the original version, the question was, “Do you change the law and make medically assisted suicide legal?” Here, answering “yes” is the CBR option, where one legalizes a form of killing (violating a moral rule) but for the benefit of more people (increasing the medical budget). Answering “no” is the rule option (follow the moral rule “you shall not kill”).
Table 1.
Example of Yes↔No reframing and Action↔Omission reframing in a vignette used in Study 2
| Original framing | Yes↔No reframing | Action↔Omission reframing |
|---|---|---|
| Do you change the law and make medically assisted suicide legal? | Do you keep the existing law where medically assisted suicide remains illegal? | Do you change the law and make medically assisted suicide illegal? |
| Yes: CBR, actionNo: Rule, inaction | Yes: Rule, inactionNo: CBR, action | Yes: Rule, actionNo: CBR, inaction |
To test for the yes–no bias, we reframed the question so that the “yes” response now corresponded to the previous “no” response (Yes↔No Reframing): “Do you keep the existing law where medically assisted suicide remains illegal?” The situation and physical action are the same as before, but which moral view corresponds to “yes” vs. “no” is swapped.
To test for omission bias, we changed whether the action coincided with the CBR option. In the original framing of “Assisted Suicide,” changing the law is the physical action and the CBR option, whereas doing nothing is the rule option. For this Action↔Omission Reframing, we created a version of the vignette where assisted suicide was currently legal, and reframed the question as, “Do you change the law and make medically assisted suicide illegal?”
This approach of reframing the same vignettes allows us to test whether responses are consistent across equivalent scenarios. It also has the advantage of ruling out any scenario-specific effects that may have been present in Study 1, as it avoids using different scenarios for action vs. omission.
For the LLMs, we used both the standard participant prompt and advice-giving prompt from Study 1. We used two new prompting techniques as a robustness check: 1) expert role prompting (58, 59), where the LLMs responded as an expert in moral philosophy, and 2) silicon sampling (60), where the LLMs generated responses from a diverse sample of synthetic subjects (based on U.S. census data and demographic information from our sample). The goal of using silicon sampling was to obtain a fair comparison between a representative sample of human participants and multiple responses from a single LLM.
In summary, we used a total of six prompts: standard participant prompt, advice-giving prompt, expert participant prompt, expert advice-giving prompt, silicon sampling with the participant prompt, and silicon sampling with the advice-giving prompt. We report the full results of the standard participant prompt in the main text. The results for the other prompts are the same in terms of the statistical significance patterns unless mentioned otherwise. Full results for the other prompts can be found in SI Appendix, section S5.
Yes–no bias.
Fig. 3 shows human and LLM responses to the reframed vignettes using the standard participant prompt. Most LLMs showed a strong bias for answering “no”; they preferred the CBR option significantly less when it coincided with “yes” than with “no” (GPT-4-turbo: 70% vs. 85%, , Llama 3.1-Instruct: 51% vs. 89%, ; Claude 3.5: 61% vs. 96%, ). GPT-4o showed a preference in the opposite direction (79% vs. 67%, ). Participants did not show a significant yes–no bias (60% vs. 56%, ).
Fig. 3.
LLMs, but not humans, show yes–no bias in Study 2. Panel (A) shows the average yes–no bias across humans and all models, and Panel (B) shows responses for each vignette. We discuss responses for the “Endowment” vignette in the main text. Error bars indicate 95% CI.
Overall, the three LLMs that preferred answering “no” were more affected by the reframing than humans (GPT-4-turbo: ; Llama 3.1-Instruct: ; Claude 3.5: ). GPT-4o, which showed a preference for answering “yes,” was also more affected by the reframing than humans, .
We found very similar results with the advice-giving prompt, the expert participant prompt, and the expert advice-giving prompt (SI Appendix, Figs. S4, S6, and S7). We again observed similar results for the participant and advice-giving prompts using silicon sampling, except we no longer found a significant effect of framing for GPT-4o (P = 0.165 and P = 0.890 respectively; SI Appendix, Figs. S10 and S11). There was a reduced ceiling effect in some individual vignettes, which suggests increased variance in responses.
Omission bias.
Fig. 4 shows responses to the reframed vignettes with the standard participant prompt. All LLMs preferred the CBR option significantly less when it coincided with action than with omission (GPT-4-turbo: 68% vs. 100%, ; GPT-4o: 83% vs. 100%, ; Llama 3.1-Instruct: 57% vs. 87%, ; Claude 3.5: 55% vs. 99%, ). Participants also showed this preference (54% vs. 66%, ).
Fig. 4.
LLMs show stronger omission bias than humans in Study 2. Panel (A) shows the average omission bias across humans and all models, and Panel (B) shows responses for each vignette. Error bars indicate 95% CI.
All models except GPT-4o showed stronger omission bias than participants (GPT-4-turbo: ; Llama 3.1-Instruct: ; Claude 3.5: ). When reframed, all LLMs flipped their preference in at least one dilemma (Fig. 4).
GPT-4o responses were not significantly different from participant responses (). However, a closer inspection of GPT-4o responses (Fig. 4B) revealed a ceiling effect in five of six vignettes, where it consistently chose the CBR option regardless of framing. The one vignette without a ceiling effect (“Endowment”) had a very strong omission bias (0% vs. 100%, ), whereas participants were less affected by the omission bias (9% vs. 59%, ). Consequently, for “Endowment,” GPT-4o showed much stronger omission bias than participants, .
With the advice-giving prompt, we found similar results except for GPT-4o responses, which no longer showed omission bias overall (75% vs. 73%), (SI Appendix, Fig. S5), and significantly less than humans, . However, GPT-4o again showed a ceiling effect in three of six vignettes. In the remaining three vignettes, it preferred omission in two and strongly preferred action in one; these responses offset each other when aggregating across dilemmas. The expert advice-giving prompt showed similar results (SI Appendix, Fig. S9).
The expert participant prompt showed similar results as the standard participant prompt (SI Appendix, Fig. S8). With silicon sampling, we again observed similar results as the main prompts but with increased variance (SI Appendix, Figs. S12 and S13).
Study 3.
Even though the moral dilemmas used in Studies 1 and 2 were based on realistic historical scenarios, they differed from the questions ordinary people might commonly ask LLMs in three key ways: 1) they contained high-stakes decisions that people rarely encounter in everyday life, 2) they were always conflicts between rules vs. CBR, and 3) the writing was more polished. To test whether our findings generalize to more naturalistic queries, we conducted a preregistered (https://osf.io/8sg4) replication with everyday dilemmas adapted from the /r/AmItheAsshole (AITA) forum on Reddit, where anonymous users share moral dilemmas they encountered to seek advice and/or feedback. In Study 3, we tested the yes–no bias and the omission bias with these naturalistic, low-stakes dilemmas, comparing the responses of 493 participants from a representative U.S. sample to those of GPT-4o, GPT-4-turbo, Llama 3.1-Instruct, and Claude 3.5.‡ The AITA dilemmas are not always conflicts between rules and CBR. Some involve a conflict between two moral rules, while others involve a conflict between what is good for oneself vs. others. Therefore, we compared responses for the original and reframed dilemmas by measuring how often people and LLMs approved of the action described in the original version of each dilemma (see Materials and Methods and Table 2 for more details).
Table 2.
Examples of Yes↔No reframing and Action↔Omission reframing in a vignette used in Study 3 (“Roommate”)
| Original framing | Yes↔No reframing | Action↔Omission reframing |
|---|---|---|
| [Vignette where you are currently in an important work meeting] Do you leave the important meeting and go help your roommate? | (Same vignette as original) Do you stay in the important meeting rather than helping your roommate? | (Vignette where you are currently with your roommate and have to attend an important work meeting soon) Do you go to your important meeting rather than helping your roommate? |
| Yes: “Original action” (roommate) | Yes: “Original omission” (meeting), inaction | Yes: “Original omission” (meeting), action |
| No: “Original omission” (meeting) | No: “Original action” (roommate), action | No: “Original action” (roommate), inaction |
We report results for the standard participant prompt in the main text. Results for the advice-giving prompt are very similar and can be found in SI Appendix, section S6.
Yes–no bias.
On average, the LLMs were again biased toward answering “no” (SI Appendix, Fig. S14; GPT-4-turbo: ; GPT-4o: ; Llama 3.1-Instruct: ; Claude 3.5: ).§ By contrast, participants did not show yes–no bias, . All LLMs were more affected by the reframing than human participants (SI Appendix, section S6).
Omission bias.
In general, LLMs and humans showed omission bias (SI Appendix, Fig. S16; Humans , GPT-4-turbo: ; GPT-4o: ; Llama 3.1-Instruct: ; Claude 3.5: ).¶
GPT-4-turbo, Llama 3.1-Instruct, and Claude 3.5 showed stronger omission bias than human participants (GPT-4-turbo: ; Llama 3.1-Instruct: , ; Claude 3.5: ). On average, the omission bias of GPT-4o was not significantly larger than that of humans, . However, inspecting all individual dilemmas without ceiling or floor effects revealed that GPT-4o either showed a significantly larger omission bias than people or a strong bias in the opposite direction (particularly in the “Pregnant” and “Roommate” vignettes; SI Appendix, Fig. S16B).
Overall, Study 3 replicated the biases found in Studies 1 and 2 under more naturalistic conditions. However, this study also revealed that there are some individual dilemmas for which some LLMs showed biases in the opposite direction (SI Appendix, Figs. S14B and S16B). This appears to be the case for those vignettes that involve self–other trade-offs and is possibly due to acquiescence (for more details, see Discussion).
Study 4.
In this study, we investigated how different methods of posttraining affect the biases of LLMs observed in Studies 1 to 3. To do so, we compared the moral decisions of different versions of Llama 3.1, namely a pretrained model and two models developed from this pretrained model through two different kinds of posttraining. One is Llama 3.1-Instruct, which was fine-tuned by Meta to “follow instructions, align with human preferences, and improve specific capabilities” (61)[p.1]. This fine-tuning consists of several rounds of reinforcement learning from human feedback (RLHF) and supervised learning from labeled examples of “good” vs. “bad” ways of responding to a curated set of queries (61). The other is Centaur, which was posttrained by cognitive science researchers on over 60,000 participants’ behavior in over 160 psychological experiments (62).
Thus, a comparison between Llama 3.1 (pretrained), Llama 3.1-Instruct, and Centaur allows us to reasonably speculate about the effects of different types of fine-tuning on the biases observed in earlier studies.
We presented Centaur and the two Llama 3.1 models with the twelve moral dilemmas used in Studies 2 and 3 (with Yes↔No Reframing and Action↔Omission Reframing) and compared their responses to the human data collected in these studies. For brevity, we only report the results for the moral dilemmas from Study 2 in the main text, but the results for dilemmas from Study 3 are extremely similar (SI Appendix, Figs. S23 and S24).
Yes–no bias.
Fig. 5A visualizes results for the yes–no bias in moral dilemmas from Study 2. Llama 3.1-Instruct showed a strong preference for “no,” , and this bias was much stronger than in the pretrained model () and Centaur (). Unlike Llama 3.1-Instruct, neither the pretrained Llama 3.1 model nor Centaur were significantly more affected by the Yes↔No Reframing than humans (SI Appendix, Table S4). Overall, these results suggest that the yes–no bias of Llama 3.1-Instruct arose from fine-tuning rather than pretraining or the architecture of the neural network.
Fig. 5.
Llama 3.1-Instruct (fine-tuned for chatbot applications) shows significantly stronger yes–no and omission bias than Llama 3.1 (Pretrained) and Centaur in Study 4. Panel (A) shows the average yes-no bias for humans and all models, and Panel (B) shows the average omission bias for humans and all models. Error bars indicate 95% CI. For responses to the individual dilemmas, see SI Appendix, Figs. S20 and S21.
Even though Centaur does not show the yes–no bias, it does not capture human responses well: Unlike humans, it shows very little variability between dilemmas, always endorsing CBR approximately 50% of the time (SI Appendix, Fig. S20B).
Omission bias.
As shown in Fig. 5B, the results for omission bias are similar to those for the yes–no bias. Llama 3.1-Instruct again showed a much stronger bias than the other models (Llama 3.1-Instruct vs. pretrained model: ; Llama 3.1-Instruct vs. Centaur: ). We found no evidence that the effect of this reframing differed between Centaur and the pretrained Llama 3.1 model, although all models’ responses differed from those of participants (SI Appendix, Table S5). Overall, this suggests that the amplified omission bias also arose from the fine-tuning Meta performed to turn their pretrained LLM into a chatbot.
Discussion
This article presents an investigation of LLM vs. human decision-making in realistic moral dilemmas and collective action problems. For moral dilemmas, commonly used LLMs showed a stronger omission bias than humans. These LLMs were biased toward choosing and advising inaction irrespective of the anticipated consequences and the imperative of the pertinent moral rules. This finding was largely robust across different prompts (experimental participant, advice-giving, role prompting, and silicon sampling) and different types of dilemmas.
Furthermore, all of these popular LLMs were sensitive to whether endorsing the choice under consideration coincided with answering “yes” or “no” (“yes–no bias”), regardless of whether the choice coincided with action vs. omission. Centaur (a LLM specifically developed to predict participants’ behavior in psychology experiments) did not exhibit these biases but also did not capture systematic differences in people’s decisions across different dilemmas. Finally, for collective action problems, LLMs gave more altruistic responses than humans. Overall, for moral dilemmas, LLM responses did not strongly correlate (all r < 0.7) with participants’ responses. For collective action problems, we only found strong correlations between participants’ responses and the responses of GPT-4-turbo and GPT-4o.
Given that omission bias is a robust phenomenon in the psychology literature (40, 41, 63), the heightened omission bias we found in LLMs is consistent with evidence that LLMs show amplified cognitive biases commonly present in human responses (e.g., by being more sensitive to experimental manipulations) (27). Further, we demonstrated that LLMs exhibit an additional bias not found in humans: the yes–no bias.
Should People Trust the Moral Advice and Moral Decisions of LLMs?.
Past research discussed whether LLMs’ moral advice is “superior” because participants rated it more favorably than the advice of other people (64) and even that of expert ethicists (11). However, the approach of using laypeople’s preferences to evaluate the quality of moral advice is problematic: Only because participants judge some moral advice more favorably does not imply that this advice is sound from the perspective of most or even any ethical theories. Further, unlike moral philosophers, LLMs are specifically trained via RLHF to provide responses that people would like, giving them an advantage in an evaluation that relies on participant ratings. In this article, we used a more objective method to assess the quality of LLMs’ moral decisions and advice: assessing whether their responses are consistent across logically equivalent questions. This method revealed that their moral advice and decisions are more biased and inconsistent than people’s.
Is it necessarily bad for LLMs to show amplified omission bias? Some ethical perspectives regard omissions as more morally permissible than actions that have the same effect [e.g., the doctrine of double effect argues that it is worse to kill than to let die; (65–67)], whereas others consider them morally equivalent [e.g., act consequentialism; (68)]. This debate is reflected in differences in how omissions are treated in different jurisdictions (69) ( 70, p.82). In some contexts, the omission bias may be unproblematic, because something being the status quo may be evidence for it working well (or mitigating downside risks). However, in many situations, including the scenarios used here, the omission bias may run counter to the greater good. Not encouraging people to act in certain situations can cause real harm to those whom the user could have helped. In such situations, certain theories of morality consider the omission immoral [e.g., utilitarianism (71)], while other moral theories do not [e.g., certain variants of deontology (72)]. The question of which moral theory is correct is beyond the scope of this article.
From a descriptive perspective, the omission bias might serve the interests of the user or the company deploying the chatbot. For instance, some users may prefer omission to avoid condemnation or punishment (73) because actions are often perceived as more causal and intentional than omissions (63, 74–76). Similarly, AI companies may prefer their chatbots to show omission bias because it might reduce their liability in jurisdictions that impose less legal liability for harms caused by inaction than for active harm (77).
The yes–no bias reflects a tendency of LLMs to provide inconsistent responses in exactly the same situation: LLMs endorse contradictory choice options depending on slight variations in the phrasing of the question. This violates an essential prerequisite for rational choice: the principle of invariance (78–81). Previous research has emphasized the importance of consistency specifically in the development of LLMs, as it is critical for ensuring that they are reliable and dependable decision-making systems (16). In Studies 2 and 3, we showed that human judgment is robust to Yes↔No Reframing, whereas most LLMs were not. Most moral philosophers would agree that when making moral decisions, one should be guided by moral principles [e.g., moral rules, social contracts, virtues, or utilitarianism (82)]. However, evidence of the yes–no bias suggests that LLMs are doing something different: They resolve moral dilemmas based on morally irrelevant, superficial differences in the wording of the question. In our studies, the inconsistency driven by the yes–no bias was present across many decisions. Although it is possible that both options may be exactly equally good in a single dilemma, it is unlikely that this was always the case across multiple scenarios as in our experiments. It is thus likely that this bias does, at least sometimes, compromise LLMs’ moral decisions and advice. Therefore, we should be reluctant to outsource our moral decisions to LLMs and critically examine the merits of their advice.
What Are Possible Sources of Biases in LLMs’ Moral Decisions?.
Different LLMs likely share features that cause systematic differences from human responses (27). Indeed, we see a similar pattern of responses for commonly used chatbot LLMs (i.e., GPT-4, Llama 3.1-Instruct, and Claude 3.5). In principle, the amplified omission bias and the new yes–no bias could arise from shared features of the network architecture, the training data, or subsequent fine-tuning and alignment efforts.
However, our results from Study 4 demonstrate that, at least in the case of Llama 3.1-Instruct, the observed biases did not arise from the network architecture or biases in the large corpus used to pretrain the model because the pretrained model did not show such strong biases. Instead, they arose from efforts to align the responses of the pretrained LLM with what the company and its users considered to be good behavior for a chatbot. For Llama 3.1-Instruct, this included multiple rounds of fine-tuning to align the responses of the pretrained Llama 3.1 model using synthetic data as well as human preference data (for details, see ref. 61, section 4). This raises the question of how the fine-tuning induced the yes–no bias even though humans do not show it. One possibility is that this bias arose through its association with omission bias. If prompts are more likely to be structured in such a way that physical action and the “yes” option correspond, the LLM might derive a tendency to answer “no” based on the confounding between the “no” answers and inaction in the scenarios it was trained on. The GPT chatbots and Claude 3.5 Sonnet were also fine-tuned using similar methods (83). While we do not have access to their pretrained base models, we speculate that the sources of their biases are likely similar.
The fine-tuning of LLMs serves multiple goals, including ensuring that the responses are harmless and ethical. Our findings suggest that while the fine-tuning involved in creating the studied chatbots might have achieved some aspects of this objective, it may have also amplified omission bias and made the model’s decisions and advice less consistent by making it highly sensitive to superficial changes in the wording of the query (i.e., the yes–no bias).
Our findings highlight a fundamental problem: The preferences and intuitions of laypeople and researchers developing these models can be a bad guide to moral AI. The fine-tuning process must be improved to ensure that LLMs make consistent and morally sound decisions. One approach would be to include multiple queries with reframed questions (as in our yes–no bias paradigm) and rewarding the model for giving consistent responses between them. However, this raises obvious issues, such as which of the different answers should be given consistently. While necessary for sound reasoning, consistent answers are not sufficient for it: The model could consistently give an answer that is morally wrong under most or all ethical frameworks. Future work on this topic will likely benefit from collaboration between AI safety researchers, moral philosophers, moral psychologists, cognitive scientists, computational ethicists, and researchers from other disciplines (84).
Limitations and Future Directions.
One limitation of our study is that, as in all survey research, our samples deviate from the general population to some extent. We mitigated this issue by using representative sampling in Studies 1 and 3 via Prolific, which has been shown repeatedly to have high data quality compared to other crowdsourcing vendors (85–87). However, while these deviations are greatly reduced for those categories that the sample was stratified across (age, sex, ethnicity, and political affiliation), there is still some deviation in others [e.g., religious identification (87)]. Further, even to the extent that the sample is representative of the US population, it may not be representative of the distribution of people who query LLMs, which likely would skew toward younger people and people who more frequently use the internet. Future work could increase representativeness by using cross-national samples and taking into account which types of people are most likely to seek advice from LLMs.
The principled method to assess the quality of LLM moral decision-making and advice developed in this article could now be applied to test future LLM versions and additional demographics with relatively little effort. It also allows researchers to test a variety of other interesting questions regarding LLM moral decision-making. Below, we propose several venues for future research that can be pursued using this method.
Further studies should systematically evaluate the logical soundness of LLMs’ decisions and advice on different moral issues, and catalog for which of those issues the responses of LLMs are logically inconsistent. Our study focused on two particular features that can distort human and LLM moral decision-making (yes vs. no framing and action vs. omission). Future research can include more morally irrelevant factors to investigate whether LLMs are also sensitive to them (e.g., the order in which information is presented (88, 89); spatial and temporal distance (90), but see ref. 91; and identifiability (92, 93), but see refs. 94 and 95). Further, it would be valuable for future work to also study the flip side of this behavior: whether LLMs are less sensitive than humans to morally relevant factors [e.g., the number of people affected (96)].
In addition to studying other moral factors, it would be interesting to further explore the boundaries and moderating conditions of the biases found here. Study 3 already points to one such boundary condition: When self–other trade-offs are concerned, the models tend to prefer answering “yes” rather than “no.” The reasons for this may be related to acquiescence (97). For instance, in a situation where someone can choose to stay somewhere or leave, asking whether one should stay indicates a preference for staying, and LLMs may try to validate the user’s opinion.# Moreover, while our paper tested a variety of prompts that gave different personas to the LLMs, another important question for future research is how the persona of the user affects LLMs’ advice. For instance, LLMs may give different advice to users depending on their social status, age, or risk tolerance. LLMs may be able to infer such information based on the user’s language. Further, ChatGPT has a memory feature that stores information about the user (98). This suggests that LLMs can take this type of user-specific information into account when answering subsequent queries.
Moreover, it would be interesting to investigate how the yes–no bias relates to other biases observed in decision-making. For instance, the preference for answering “no” is reminiscent of the effect of inclusion vs. exclusion framing, where participants are more restrictive when asked to include than when asked to exclude (99, 100). This effect has also been documented in the moral domain, where an exclusion framing leads to a larger moral circle compared to inclusion framing (101). Alternatively, it could be linked to the default effect (102), assuming that “no” is the default answer that the model gives when it cannot decide. Future research may investigate to what extent LLMs show these related biases.
Another vital topic for future research is to what extent people follow the advice they receive from LLMs. Prior research on advice-taking shows that people often discount advice, particularly when it is unsolicited, from a novice, or conflicts with their prior beliefs (103–105). It is an open question how much people trust advice from LLMs compared to human advisors. While there is some evidence that LLMs’ moral advice can influence people’s decisions (10), how much people trust this advice could depend on whether they consider the LLMs to be experts in morality (11, 64) and how much it conflicts with their prior beliefs. Due to their training on human data and human feedback, LLMs likely tend to give responses that people would like. Given the prevalence of confirmation bias (106–108), this may increase users’ reliance on their advice. This could be problematic: As we demonstrate in this paper, these responses, even though people might like them, may contain new biases or amplify existing human biases.
Finally, while the results from Study 4 suggest that the biases arose from fine-tuning the pretrained LLMs into chatbots, it remains unclear which component(s) of the fine-tuning process caused these biases (see also ref. 61, section 4). Investigating how AI companies fine-tune chatbots and analyzing how different elements of the fine-tuning process affect the biases studied in this article is an important direction for future research. This could be studied by creating a large set of models, each fine-tuned with different elements of Llama 3.1-Instruct’s fine-tuning process and investigating which of them show amplified biases.
Conclusion.
The moral decision-making of LLMs is biased and has significant room for improvement. Characterizing, understanding, and overcoming these limitations should be a key priority for future work on LLMs. Study 4 suggested that the observed biases are not inherent in the network architecture or the text corpora that the LLMs are trained on. Instead, they seem to result from how AI companies fine-tune LLMs to develop them into chatbots that adhere to the companies’ rules and produce responses that are desirable to their consumers. If this is the case, then it might be relatively easy for AI companies to rectify the biases documented in this article by making adjustments to the fine-tuning process of existing models.
To further characterize the nature and limitations of LLMs’ moral reasoning, they should be evaluated on additional tests [e.g., the defining issues test (109)]. Accurately assessing their capacity for moral reasoning will require a large battery of novel automated benchmarks that are valid and reliable.
We hope that our research and also other research in this field will inform future improvements in the moral decisions and advice of LLMs. Hopefully, this can inform laws and policies for mitigating risks from advanced AI by prohibiting morally irresponsible applications of LLMs and incentivizing the development of safe and ethical AI.
Materials and Methods
All experiments received ethical approval from the Office of the Human Research Protection Program, The University of California, Los Angeles (UCLA OHRPP) under protocol number IRB#23-001436. Informed consent was obtained from all participants. This article contains SI Appendix online at (TBA).
Study 1.
Participants.
We recruited 294 participants‖ from a representative U.S. sample on Prolific, which is based on US census data from 2021 and stratified across age, sex, ethnicity, and political affiliation.** Participants were paid $4 for the 25-min study (base rate of $3.33 and a bonus of $0.67 if they passed all attention checks). Nine participants failed one of the two attention checks asking about details of a dilemma they had been shown on the previous page, leaving us with a final sample of 285 participants. The mean age was 45.53 (); 146 participants were male, and 139 female; 28 participants identified as Asian, 37 as Black, 33 as Mixed, 159 as White, and 28 as other; 88 participants identified as Democrats, 117 as Independents, and 80 as Republicans.
We used four commonly used LLMs: GPT-4o, GPT-4-turbo (110), Llama 3.1-Instruct 70B (111), and Claude 3.5 Sonnet (for details, see SI Appendix, section S1).
Design and materials.
Moral dilemmas.
For the moral dilemmas, we used a set of 13 vignettes from ref. 31, which were originally developed by ref. 42 and adapted for clarity and removing potential confounds (e.g., by making it clear that the decision-maker is not impacted by the outcomes of their decisions). The full set of materials is available at: https://osf.io/ybdr9.
To eliminate the confound between action with the CBR option and omission with the rule option, the framing for the choice action was varied across vignettes (SI Appendix, section S2.1). In eight vignettes, the CBR option coincided with action (Action Framing), and in five vignettes, it coincided with omission (Omission Framing).
Collective action problems.
For the collective action problems, we used a set of nine vignettes from refs. 48 and 49 (SI Appendix, section S2.2), where the conflict is between self-interest and the greater good. Most problems were framed so that higher values on the slider indicate more altruistic decisions (as in the example, switching more hours to volunteering is more altruistic), and only reversed in one vignette (“Drought”).
System prompts.
We used two different system prompts (experimental participant vs. advice-giving) before showing a dilemma to the LLMs. The experimental participant prompt was designed to mimic what researchers might do to simulate a psychology experiment with LLMs. The advice-giving prompt was designed to mimic how people ask a chatbot for advice. For the advice-seeking version of the collective action problem, we rephrased the dilemma to a first-person perspective, where the user asks for advice. All prompts are included in the SI Appendix, section S3.
Procedure.
Participants recruited on Prolific completed an online study on Qualtrics, where they read all 13 moral dilemmas and all 9 collective action problems. Participants were randomly assigned to either read the moral dilemmas first or the collective action problems first. The order of vignettes within each type of dilemma was randomized. They read and made a decision for each vignette.
For the LLMs, we showed the prompt followed by the vignette. Each LLM responded only to a single dilemma at a time, after which a new session was created to query the next dilemma. The reason for querying the LLMs in this way, rather than asking all dilemmas sequentially, was to keep it more similar to how LLMs would typically be queried by a user (who would rarely ask about a sequence of 22 dilemmas in a single session).
Data Analysis.
For the effect of framing on LLMs, we used a linear regression analysis. For the effect of framing on participants, we used a linear mixed effects model to account for the fact that the same participant responded to multiple dilemmas (see ref. 112 for a discussion on the benefits of applying linear models to binary data. We also conducted a robustness check using logistic models, which lead to very similar results). We used the afex package (113) in R with effect contrast coding, which is the default in this package.
To compare mean altruism between participants and LLMs for the collective action problems, we first took the mean across all nine dilemmas for each participant and for each set of LLM responses to each dilemma (i.e., we average across dilemmas and obtained 500 averages per LLM, equivalent to the number of answers per dilemma). We then t tested those means between models (rather than testing the data points directly to not overweight the dependent responses for participants, where each participant answered nine dilemmas).
For correlations between LLM and human responses, we first aggregated the data within dilemmas, then calculated a correlation of within-dilemma means for different models and prompts.
Study 2.
We preregistered Study 2 at: https://osf.io/t4w9g.
Participants.
We recruited 501 participants and excluded 11 based on a preregistered attention check, which asked what the scenario was about, leaving us with a final sample of N = 474. The mean age was 40.11 (SD = 13.20); 237 participants were female, 236 male, and one did not share their gender. Participants were paid $0.48 for the 3-min study.
Large language models.
We used the same LLMs and parameters as in Study 1, except that in this study, we had preregistered using Llama 3-Instruct (which we refer to as Llama 3) and Claude 3 Opus (which we refer to as Claude 3) before updating to the more recent Llama 3.1-Instruct and Claude 3.5 Sonnet. Therefore, we also present results for Llama 3 and Claude 3 for the standard participant prompt in the SI Appendix, section S5.
Design and materials.
This study used a between-subjects design, where participants were randomly assigned to read and make a decision on one vignette. LLMs also saw only one vignette with each query. We used a subset of six moral dilemma vignettes from Study 1 where the action under consideration could be reasonably reframed. In addition to the original framing version of each dilemma, we showed participants two reframed versions of the dilemma, Yes↔No Reframing and Action↔Omission Reframing (Table 1; full materials at: https://osf.io/ybdr9). Participants randomly saw one of 18 possible vignettes (one of three framing versions of the six moral dilemmas).
For Yes↔No Reframing, we reframed the vignette such that the question reversed the “yes” or “no” responses referred to a given choice (compared to the original framing). For Action↔Omission Reframing, we rewrote the vignette and reframed the question such that the response that had corresponded with action now corresponded with omission, and vice versa. To measure the extent to which responses are afflicted by each bias, we compared responses of dilemmas with each type of reframing to the original framing of the dilemmas.
System prompts.
We used the standard participant prompt and advice-giving prompt from Study 1. We also added a new prompt where we asked the LLM to respond as if it was an expert in moral philosophy. This is based on the prompt-engineering technique “role-prompting” (58, 59) (SI Appendix, section S3.5).
We also employed silicon sampling (60) to simulate responses from a diverse human sample. We gave the LLMs a revised prompt in which they simulated “silicon” individuals by randomizing demographic characteristics such as age, gender, ethnicity, socioeconomic status, education level, and political and religious affiliation from demographic information from our sample or, when not available, from U.S. census data. We selected these demographic characteristics both for the purpose of simulating diversity and also based on extant literature demonstrating their relationship with moral decision-making (e.g., religiosity, ref. 114; political affiliation and other demographic characteristics, ref. 115). We generated a text template with these different demographic characteristics as template fragments. Then, we randomized the template fragments to create a diverse “silicon” sample to be used as part of a prompt for the LLMs (SI Appendix, section S3.6). We included this text after the standard participant prompt.
Data analysis.
We tested each framing effect using an ANOVA with the main effects of framing, model (each of the LLMs and humans), and vignette, and interactions between these factors.
We preregistered testing the following hypotheses: For Yes↔No Reframing, we predicted that the LLMs would show a systematic bias to answering “no” when making decisions, that humans would not show this bias, and that this bias would consequently be larger for LLMs than for humans. For Action↔Omission Reframing, we predicted that both LLMs and humans would show omission bias, and this bias would be larger for LLMs than for humans.
Study 3.
We preregistered Study 3 at: https://osf.io/8sg4p.
Participants.
We recruited 497 participants†† from a representative U.S. sample on Prolific (for details about representative sampling, see Study 1). We excluded four participants based on a preregistered attention check. Our final sample was N = 491. The mean age was 45.5 (SD = 15.8); 251 participants were female and 240 male; 36 participants identified as Asian, 56 as Black, 50 as Mixed, 36 as Other, and 313 as White; 146 participants identified as Democrats, 211 as Independents, and 134 as Republicans. Participants were paid $0.64 for the 4-min study.
Large language models.
We used the same LLMs and parameters as in Study 2, except we preregistered using Claude 3.5 Sonnet instead of Claude 3 Opus. We had originally preregistered using Llama 3 before updating to the more recent Llama 3.1-Instruct model. Therefore, we also present results for Llama 3 for the preregistered standard participant prompt in the SI Appendix, section S6.
Design and materials.
We used moral dilemmas from an online forum (AITA on Reddit, https://www.reddit.com/r/AmItheAsshole). We used six posts from a large dataset of AITA posts (116) to develop a new set of vignettes that are less polished and have lower stakes compared to the moral dilemmas in Studies 1 and 2 (SI Appendix, section S2.1; full materials at: https://osf.io/3ahrb). As in Study 2, in addition to the dilemma with original framing, we adapted the vignettes to include Yes↔No Reframing and Action↔Omission Reframing versions of the dilemma. Participants randomly saw one of 18 possible vignettes (one of three framing versions of the six moral dilemmas; Table 2).
We selected the posts by randomly subsetting the data and reading through the posts, choosing ones that were appropriate (e.g., did not include sensitive topics), could be reasonably reframed, and constituted a moral dilemma. We then rewrote these posts in the second person and the present tense (instead of the original past-tense, first-person narrative) and removed some irrelevant or overly emotive details. The purpose was to make these vignettes more consistent with those used in previous studies and commonly seen moral dilemmas and to reduce demand characteristics by ensuring that the vignette would not be interpreted as a narrator seeking validation for something they had already done. We piloted and revised the dilemmas to ensure that each was balanced (i.e., that participants would not overwhelmingly choose one option over another). We kept most of the phrasing used in the original posts to keep them naturalistic.
Many of these dilemmas do not necessarily contrast a CBR option with a rule option like the ones used previously. Some could be interpreted as contrasting two moral rules (e.g., obligation to a friend vs. work in “Roommate”), or helping others vs. self-interest (e.g., staying at home with your wife who is eight months pregnant vs. enjoying game night with your friends). Further, the decision-maker would always expect to be affected by the outcomes of their decisions.
System prompts.
We shortened the system prompt to make it consistent with the instructions that participants saw (SI Appendix, sections S3.7 and S3.8).
Data analysis.
We preregistered and conducted the same data analysis and hypotheses as in Study 2.
Study 4.
We used the human participant data we collected from Studies 2 and 3. For the LLMs (Centaur, Llama 3.1-Instruct, and the pretrained Llama 3.1 model), we only used the participant prompt, as we were interested in how the models can approximate human participant responses. The pretrained Llama model was not designed for advice-giving, and neither was Centaur, which was designed to simulate participant responses in psychology experiments rather than to give advice (62). More details on prompting can be found in SI Appendix, section S1.
Supplementary Material
Appendix 01 (PDF)
Acknowledgments
This work was partially supported by a grant from Forethought Foundation for Global Priorities Research to M.M. We thank Gilad Feldman, Mengxuan Helen Qiao, and the Causal Cognition Lab at University College London for helpful suggestions and Marcel Binz for guidance on querying the Centaur model.
Author contributions
V.C., M.M., and F.L. designed research; V.C. and M.M. performed research; V.C. and M.M. analyzed data; V.C. and M.M. checked code; F.L. supervised the project; and V.C., M.M., and F.L. wrote the paper.
Competing interests
The authors declare no competing interest.
Footnotes
This article is a PNAS Direct Submission.
*We did not run a between-subjects study because, at the time, the cost of recruiting a representative sample on Prolific increased with the number of participants (independent of the duration of the study). Therefore, a between-subjects design where participants only see one dilemma would have increased the cost by a substantial amount. In Studies 2 and 3, we use a between-subjects design to rule out any influence of within vs. between-subjects manipulation.
†We had originally preregistered to use Claude 3 Opus and Llama 3 because the more recent Claude 3.5 Sonnet and Llama 3.1-Instruct were not available at the time. We report the results for the most recent models in the main text. Results for Claude 3 Opus and Llama 3 can be found in the SI Appendix, Figs. S2 and S3).
‡We originally preregistered using Llama 3, but later updated to Llama 3.1, the most capable Llama model at the time of writing. Results for Llama 3 can be found in SI Appendix, Figs. S18 and S19.
§Llama 3 showed a significant bias toward answering “yes” (SI Appendix, Fig. S18).
¶Llama 3 showed a significant action bias (SI Appendix, Fig. S19).
#This was not an issue for the collective action problems in Study 1, where the question was framed in a neutral way (e.g., “How much do you allocate to...”) rather than in a way where a specific choice option would correspond to “yes.”
‖We had originally intended to recruit 300 participants; however, only 294 completed the survey in a reasonable time frame.
**More information about what census data and allocation algorithm is used by Prolific can be found at: https://researcher-help.prolific.com/en/article/e6555f.
††We had originally intended to recruit 500 participants; however, only 497 completed the survey in a reasonable time frame.
Data, Materials, and Software Availability
Anonymized data and materials data have been deposited in OSF (https://osf.io/3kvjd/) (117).
Supporting Information
References
- 1.M. Fraiwan, N. Khasawneh, A review of ChatGPT applications in education, marketing, software engineering, and healthcare: Benefits, drawbacks and research directions. arXiv [Preprint] (2023). 10.48550/arXiv.2305.00237 (Accessed 26 June 2024). [DOI]
- 2.Messeri L., Crockett M., Artificial intelligence and illusions of understanding in scientific research. Nature 627, 49–58 (2024). [DOI] [PubMed] [Google Scholar]
- 3.OpenAI, Introducing the model spec: Transparency in OpenAI’s models. https://openai.com/index/introducing-the-model-spec/. Accessed 10 May 2024.
- 4.Y. Zhang et al. , Siren’s song in the AI ocean: A survey on hallucination in large language models. arXiv [Preprint] (2023). 10.48550/arXiv.2309.01219. [DOI]
- 5.J. Echterhoff, Y. Liu, A. Alessa, J. McAuley, Z He, Cognitive bias in high-stakes decision-making with LLMs. arXiv [Preprint] (2024). 10.48550/arXiv.2403.00811 (Accessed 26 June 2024). [DOI]
- 6.Y. Gao et al. , Chat with ChatGPT on interactive engines for intelligent driving. IEEE Trans. Intell. Veh. 8, 2034–2036 (2023).
- 7.Takemoto K., The moral machine experiment on large language models. R. Soc. Open Sci. 11, 231393 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Lei L., Zhang H., Yang S. X., ChatGPT in connected and autonomous vehicles: Benefits and challenges. Intell. Robot 3, 145–148 (2023). [Google Scholar]
- 9.Joo Y. K., Kim B., Selfish but socially approved: The effects of perceived collision algorithms and social approval on attitudes toward autonomous vehicles. Int. J. Human Comput. Inter. 39, 3717–3727 (2023), 10.1080/10447318.2022.2102716. [DOI] [Google Scholar]
- 10.Krügel S., Ostermaier A., Uhl M., ChatGPT’s inconsistent moral advice influences users’ judgment. Sci. Rep. 13, 4569 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.D. Dillion, D. Mondal, N. Tandon, K. Gray, AI language model rivals expert ethicist in perceived moral expertise. Sci. Rep. 15, 4084 (2025), 10.1038/s41598-025-86510-0. [DOI] [PMC free article] [PubMed]
- 12.Liang W., et al. , Advances, challenges and opportunities in creating data for trustworthy AI. Nat. Mach. Intell. 4, 669–677 (2022). [Google Scholar]
- 13.Kaplan S., Handelman D., Handelman A., Sensitivity of neural networks to corruption of image classification. AI Ethics 1, 1–10 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.N. Shapira et al. , Clever hans or neural theory of mind? Stress testing social reasoning in large language models. arXiv [Preprint] (2023). 10.48550/arXiv.2305.14763 (Accessed 26 June 2024). [DOI]
- 15.M Sap, R LeBras, D Fried, Y Choi, Neural theory-of-mind? on the limits of social intelligence in large LMs. arXiv [Preprint] (2022). 10.48550/arXiv.2210.13312 (Accessed 26 June 2024). [DOI]
- 16.Y. Liu et al. , Aligning with logic: Measuring, evaluating and improving logical consistency in large language models. arXiv [Preprint] (2024). http://arxiv.org/abs/2410.02205 (Accessed 10 March 2025).
- 17.T. Räuker, A. Ho, S. Casper, D. Hadfield-Menell, “Toward transparent AI: A survey on interpreting the inner structures of deep neural networks” in 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) (Raleigh, NC, 2023), pp. 464-483, 10.1109/SaTML54575.2023.00039. [DOI]
- 18.Y. Bengio et al. , Managing AI risks in an era of rapid progress. arXiv [Preprint] (2023). 10.48550/arXiv.2310.17688 (Accessed 26 June 2024). [DOI]
- 19.Binz M., Schulz E., Using cognitive psychology to understand GPT-3. Proc. Natl. Acad. Sci. U.S.A. 120, e2218523120 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Hagendorff T., Fabi S., Kosinski M., Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT. Nat. Comput. Sci. 3, 833–838 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Park P. S., Schoenegger P., Zhu C., Diminished diversity-of-thought in a standard large language model. Behav. Res. Methods 56, 5754–5770 (2024), 10.3758/s13428-023-02307-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.J. Coda-Forno, M. Binz, J. X. Wang, E. Schulz, Cogbench: A large language model walks into a psychology lab. arXiv [Preprint] (2024). 10.48550/arXiv.2402.18225 (Accessed 5 October 2024). [DOI]
- 23.Gawronski B., Beer J. S., What makes moral dilemma judgments "utilitarian" or "deontological"? Soc. Neurosci. 12, 626–632 (2017). [DOI] [PubMed] [Google Scholar]
- 24.P. Foot., The problem of abortion and the doctrine of the double effect. Oxford Rev. 5, 5–15 (1967).
- 25.Thomson J. J., Killing, letting die, and the trolley problem. Monist 59, 204–217 (1976). [DOI] [PubMed] [Google Scholar]
- 26.Rehman U., Iqbal F., Shah M. U., Exploring differences in ethical decision-making processes between humans and ChatGPT-3 model: a study of trade-offs. AI Ethics 5, 279–289 (2023), 10.1007/s43681-023-00335-z. [DOI] [Google Scholar]
- 27.Almeida G. F., Nunes J. L., Engelmann N., Wiegmann A., de Araújo M., Exploring the psychology of LLMs’ moral and legal reasoning. Artif. Intell. 333, 104145 (2024). [Google Scholar]
- 28.Bauman C. W., McGraw A. P., Bartels D. M., Warren C., Revisiting external validity: Concerns about trolley problems and other sacrificial dilemmas in moral psychology. Soc. Pers. Psychol. Compass 8, 536–554 (2014). [Google Scholar]
- 29.Bennis W. M., Medin D. L., Bartels D. M., The costs and benefits of calculation and moral rules. Perspect. Psychol. Sci. 5, 187–202 (2010). [DOI] [PubMed] [Google Scholar]
- 30.Kahane G., Sidetracked by trolleys: Why sacrificial moral dilemmas tell us little (or nothing) about utilitarian judgment. Soc. Neurosci. 10, 551–560 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.M. Maier, V. Cheung, F. Lieder, Learning from outcomes shapes reliance on moral rules versus cost-benefit reasoning. PsyArXiv [Preprints] (2025). https://osf.io/preprints/psyarxiv/gjf3h_v4 (Accessed 26 June 2024).
- 32.M. Maier, V. Cheung, F. Lieder, A reinforcement-learning meta-control architecture based on the dual-process theory of moral decision-making. PsyArXiv [Preprints] (2023). 10.31234/osf.io/j6fhk (Accessed 26 June 2024). [DOI]
- 33.G. Gigerenzer, “Moral intuition = fast and frugal heuristics?” in Moral Psychology, W. Sinnott-Armstrong, Ed. (MIT Press, 2008), pp. 1–26.
- 34.Harsanyi J. C., Rule utilitarianism and decision theory. Erkenntnis 11, 25–53 (1977). [Google Scholar]
- 35.Crone D. L., Laham S. M., Utilitarian preferences or action preferences? De-confounding action and moral code in sacrificial dilemmas Pers. Individ. Differ. 104, 476–481 (2017). [Google Scholar]
- 36.Körner A., Deutsch R., Gawronski B., Using the CNI model to investigate individual differences in moral dilemma judgments. Pers. Soc. Psychol. Bull. 46, 1392–1407 (2020). [DOI] [PubMed] [Google Scholar]
- 37.Gawronski B., Armstrong J., Conway P., Friesdorf R., Hütter M., Consequences, norms, and generalized inaction in moral dilemmas: The CNI model of moral decision-making. J. Pers. Soc. Psychol. 113, 343–376 (2017). [DOI] [PubMed] [Google Scholar]
- 38.Ritov I., Baron J., Reluctance to vaccinate: Omission bias and ambiguity. J. Behav. Decis. Mak. 3, 263–277 (1990). [Google Scholar]
- 39.Kahneman D., Tversky A., The psychology of preferences. Sci. Am. 246, 160–173 (1982), 10.1038/scientificamerican0182-160. [DOI] [Google Scholar]
- 40.Baron J., Ritov I., Omission bias, individual differences, and normality. Organ. Behav. Hum. Decis. Process. 94, 74–85 (2004). [Google Scholar]
- 41.Yeung S. K., Yay T., Feldman G., Action and inaction in moral judgments and decisions: Meta-analysis of omission bias omission-commission asymmetries. Pers. Soc. Psychol. Bull. 48, 1499–1515 (2022), 10.1177/01461672211042315. [DOI] [PubMed] [Google Scholar]
- 42.Körner A., Deutsch R., Deontology and utilitarianism in real life: A set of moral dilemmas based on historic events. Pers. Soc. Psychol. Bull. 49, 1511–1528 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Dawes R. M., Social dilemmas. Annu. Rev. Psychol. 31, 169–193 (1980). [Google Scholar]
- 44.Van Lange P. A., Joireman J., Parks C. D., Van Dijk E., The psychology of social dilemmas: A review. Organ. Behav. Hum. Decis. Process. 120, 125–141 (2013). [Google Scholar]
- 45.C. D. Parks, Trust and Social Dilemmas (Oxford Research Encyclopedia of Psychology, 2019). 10.1093/acrefore/9780190236557.013.443. [DOI]
- 46.R. Suleiman, D. V. Budescu, I. Fischer, D. M. Messick, Eds., Contemporary Psychological Research on Social Dilemmas (Cambridge University Press, 2004).
- 47.Hardin G., The tragedy of the commons. Science 162, 1243–1248 (1968). [PubMed] [Google Scholar]
- 48.T. Burga et al. , “Decision-makers systematically overlook crucial consideration in social dilemmas: implications for finding public policies that maximize collective well-being” in Well-Being, Public Policies, and Sustainable Human Development (Paris, France, 2023), 10.13140/RG.2.2.10814.05449/2. [DOI]
- 49.P. Groß et al. , “What (doesn’t) limit people’s prosociality in social dilemma situations” in 11th European Conference on Positive Psychology (Innsbruck, Austria, 2024), 10.13140/RG.2.2.10049.12649/1. [DOI]
- 50.Grossmann I., et al. , AI and the transformation of social science research. Science 380, 1108–1109 (2023). [DOI] [PubMed] [Google Scholar]
- 51.Dillion D., Tandon N., Gu Y., Gray K., Can AI language models replace human participants? Trends Cogn. Sci. 27, 597–600 (2023). [DOI] [PubMed] [Google Scholar]
- 52.Cao X., Kosinski M., Large language models know how the personality of public figures is perceived by the general public. Sci. Rep. 14, 6735 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.M. Miotto, N. Rossberg, B. Kleinberg, Who is GPT-3? An exploration of personality values and demographics. arXiv [Preprint] (2022). 10.48550/arXiv.2209.14338 (Accessed 26 June 2024). [DOI]
- 54.Ouyang L., et al. , Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 35, 27730–27744 (2022). [Google Scholar]
- 55.W. Agnew et al. , “The illusion of artificial inclusion” in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Association for Computing Machinery, New York, NY, 2024).
- 56.M. Atari, M. J. Xue, P. S. Park, D. Blasi, J. Henrich (2023). Which humans? PsyArXiv [Preprints] (2023). 10.31234/osf.io/5b26t (Accessed 26 June 2024). [DOI]
- 57.A. Wang, J. Morgenstern, J. P. Dickerson, Large language models cannot replace human participants because they cannot portray identity groups. arXiv [Preprint] (2024). 10.48550/arXiv.2402.01908 (Accessed 26 June 2024). [DOI]
- 58.B. Chen, Z. Zhang, N. Langrené, S. Zhu, Unleashing the potential of prompt engineering in large language models: A comprehensive review. arXiv [Preprint] (2023). 10.48550/arXiv.2310.14735 (Accessed 5 October 2024). [DOI]
- 59.S. Ekin, Prompt engineering for ChatGPT: a quick guide to techniques, tips, and best practices. TechRxiv [Preprints] (2023). 10.36227/techrxiv.22683919.v2 (Accessed 5 October 2024). [DOI]
- 60.Argyle L. P., et al. , Out of one, many: Using language models to simulate human samples. Polit. Anal. 31, 337–351 (2023). [Google Scholar]
- 61.A. Dubey et al. , The Llama 3 herd of models. arXiv [Preprint] (2024). 10.48550/arXiv.2407.21783 (Accessed 5 October 2024). [DOI]
- 62.M. Binz et al. , Centaur: A foundation model of human cognition. PsyArXiv [Preprints] (2024). 10.31234/osf.io/d6jeb (Accessed 5 Oct 2024). [DOI]
- 63.Feldman G., Kutscher L., Yay T., Omission and commission in judgment and decision making: Understanding and linking action-inaction effects using the concept of normality. Soc. Pers. Psychol. Compass 14, e12557 (2020), 10.1111/spc3.12557. [DOI] [Google Scholar]
- 64.Aharoni E., et al. , Attributions toward artificial agents in a modified Moral Turing Test. Sci. Rep. 14, 8458 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Persson I., The act-omission doctrine and negative rights. J. Value Inq. 41, 15 (2007). [Google Scholar]
- 66.H. Spector, 187the moral asymmetry between acts and omissions in Legal, Moral, and Metaphysical Truths: The Philosophy of Michael S. Moore (Oxford University Press, 2016), 10.1093/acprof:oso/9780198703242.003.0013. [DOI]
- 67.Quinn W. S., Actions, intentions, and consequences: The doctrine of double effect. Philos. Public Aff. 18, 334–351 (1989). [PubMed] [Google Scholar]
- 68.Baron J., Morality and Rational Choice (Vol. 18) (Springer Science & Business Media, 1993). [Google Scholar]
- 69.Nelkin D. K., Rickless S. C., The Ethics and Law of Omissions (Oxford University Press, 2017). [Google Scholar]
- 70.Glendon M. A., Rights Talk: The Impoverishment of Political Discourse (Simon and Schuster, 2008). [Google Scholar]
- 71.T. John, “Mozi [Accessed: 2024-06-26]” in Introduction to Utilitarianism, R. Y. Chappell, D. Meissner, W. MacAskill, Eds. (2023). https://www.utilitarianism.net/utilitarianthinker/mozi.
- 72.Alexander L., Moore M., “Deontological ethics” in The Stanford Encyclopedia of Philosophy (Winter 2021), E. N. Zalta, Ed. (Metaphysics Research Lab, Stanford University, 2021).
- 73.DeScioli P., Christner J., Kurzban R., The omission strategy. Psychol. Sci. 22, 442–446 (2011). [DOI] [PubMed] [Google Scholar]
- 74.Jamison J., Yay T., Feldman G., Action-inaction asymmetries in moral scenarios: Replication of the omission bias examining morality and blame with extensions linking to causality, intent, and regret. J. Exp. Soc. Psychol. 89, 103977 (2020). [Google Scholar]
- 75.Feltz A., May J., The means/side-effect distinction in moral cognition: A meta-analysis. Cognition 166, 314–327 (2017). [DOI] [PubMed] [Google Scholar]
- 76.Spranca M., Minsk E., Baron J., Omission and commission in judgment and choice. J. Exp. Soc. Psychol. 27, 76–105 (1991), 10.1016/0022-1031(91)90011-T. [DOI] [Google Scholar]
- 77.K. K. Ferzan, “Omissions, acts, and the duty to rescue” in The Ethics and Law of Omissions, D. K. Nelkin, S. C. Rickless, Eds. (Oxford University Press, 2017), pp. 217–234.
- 78.J. Von Neumann, O. Morgenstern, Theory of Games and Economic Behavior (Princeton University Press, 2nd rev. Ed., 1947).
- 79.A. Tversky, D. Kahneman, “Rational choice and the framing of decisions” in Decision Making: Descriptive, Normative, and Prescriptive Interactions, D. E. Bell, H. Raiffa, A. Tversky, Eds. (Cambridge University Press, 1988), pp. 167–192.
- 80.Schick F., Consistency and rationality. J. philos. 60, 5–19 (1963). [Google Scholar]
- 81.Sugden R., Why be consistent? A critical analysis of consistency requirements in choice theory Economica 52, 167–183 (1985). [Google Scholar]
- 82.Parfit D., On What Matters (Oxford University Press, 2011). [Google Scholar]
- 83.Y. Wang et al. , Aligning large language models with human: A survey. arXiv [Preprint] (2023). 10.48550/arXiv.2307.12966 (Accessed 5 October 2024). [DOI]
- 84.Awad E., et al. , Computational ethics. Trends Cogn. Sci. 26, 388–405 (2022). [DOI] [PubMed] [Google Scholar]
- 85.Douglas B. D., Ewell P. J., Brauer M., Data quality in online human-subjects research: Comparisons between mturk, prolific, cloudresearch, qualtrics, and sona. PLoS One 18, e0279720 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Eyal P., David R., Andrew G., Zak E., Ekaterina D., Data quality of platforms and panels for online behavioral research. Behav. Res. Methods 54, 1–20 (2021).34085234 [Google Scholar]
- 87.M. N. Stagnaro et al. , Representativeness versus attentiveness: A comparison across nine online survey samples. PsyArXiv [Preprint] (2024). 10.31234/osf.io/h9j2d (Accessed 5 Oct 2024). [DOI]
- 88.Wiegmann A., Okan Y., Nagel J., Order effects in moral judgment. Philos. Psychol. 25, 813–836 (2012). [Google Scholar]
- 89.Schwitzgebel E., Cushman F., Expertise in moral reasoning? Order effects on moral judgment in professional philosophers and non-philosophers Mind Lang. 27, 135–153 (2012), 10.1111/j.1468-0017.2012.01438.x. [DOI] [Google Scholar]
- 90.Trope Y., Liberman N., Construal level theory. Handb. Theor. Soc. Psychol. 1, 118–134 (2012). [Google Scholar]
- 91.M. Maier et al. , Adjusting for publication bias reveals that evidence for and size of construal level theory effects is substantially overestimated. PsyArXiv [Preprint] (2022). 10.31234/osf.io/r8nyu (Accessed 5 Oct 2024). [DOI]
- 92.Kogut T., Ritov I., The “identified victim” effect: An identified group, or just a single individual? J. Behav. Decis. Mak. 18, 157–167 (2005). [Google Scholar]
- 93.Small D. A., Loewenstein G., Slovic P., Sympathy and callousness: The impact of deliberative thought on donations to identifiable and statistical victims. Organ. Behav. Hum. Decis. Process. 102, 143–153 (2007), 10.1016/j.obhdp.2006.01.005. [DOI] [Google Scholar]
- 94.R. Majumder, Y. L. C. Tai, I. Ziano, G. Feldman, Revisiting the impact of singularity on the identified victim effect: Replication and extension of Kogut and Ritov (2005a) Study 2. (2024). Open Science Framework. 10.17605/OSF.IO/9QCPJ. Accessed 5 Oct 2024. [DOI]
- 95.M. Maier, Y. C. Wong, G. Feldman, Revisiting and rethinking the identifiable victim effect: Replication and extension of Small, Loewenstein, and Slovic (2007). Collabra Psychol. 9, 90203 (2023), 10.1525/collabra.90203. [DOI]
- 96.W. H. Desvousges et al. , “Measuring natural resource damages with contingent valuation: Tests of validity and reliability” in Contributions to Economic Analysis, B. H. Baltagi, F. Moscore, Eds. (Emerald Group Publishing Limited, 1993), pp. 91–164. 10.1016/B978-0-444-81469-2.50009-2. [DOI]
- 97.Tjuatja L., Chen V., Wu T., Talwalkwar A., Neubig G., Do LLMs exhibit human-like response biases? A case study in survey design Trans. Assoc. Comput. Linguist. 12, 1011–1026 (2024), 10.1162/tacla00685. [DOI] [Google Scholar]
- 98.OpenAI. Memory and new controls for ChatGPT. https://openai.com/index/memory-and-new-controls-for-chatgpt/. Accessed 8 December 2024.
- 99.Yaniv I., Schul Y., Elimination and inclusion procedures in judgment. J. Behav. Decis. Mak. 10, 211–220 (1997). [Google Scholar]
- 100.Yaniv I., Schul Y., Raphaelli-Hirsch R., Maoz I., Inclusive and exclusive modes of thinking: Studies of prediction, preference, and social perception during parliamentary elections. J. Exp. Soc. Psychol. 38, 352–367 (2002). [Google Scholar]
- 101.Laham S. M., Expanding the moral circle: Inclusion and exclusion mindsets and the circle of moral regard. J. Exp. Soc. Psychol. 45, 250–253 (2009). [Google Scholar]
- 102.Jachimowicz J. M., Duncan S., Weber E. U., Johnson E. J., When and why defaults influence decisions: A meta-analysis of default effects. Behav. Public Policy 3, 159–186 (2019). [Google Scholar]
- 103.Bonaccio S., Dalal R. S., Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences. Organ. Behav. Hum. Decis. Process. 101, 127–151 (2006). [Google Scholar]
- 104.Yaniv I., Receiving other people’s advice: Influence and benefit. Organ. Behav. Hum. Decis. Process. 93, 1–13 (2004). [Google Scholar]
- 105.Harvey N., Fischer I., Taking advice: Accepting help, improving judgment, and sharing responsibility. Organ. Behav. Hum. Decis. Process. 70, 117–133 (1997). [Google Scholar]
- 106.Hergovich A., Schott R., Burger C., Biased evaluation of abstracts depending on topic and conclusion: Further evidence of a confirmation bias within scientific psychology. Curr. Psychol. 29, 188–209 (2010). [Google Scholar]
- 107.Masnick A. M., Zimmerman C., Evaluating scientific research in the context of prior belief: Hindsight bias or confirmation bias. J. Psychol. Sci. Technol. 2, 29–36 (2009). [Google Scholar]
- 108.Nickerson R. S., Confirmation bias: A ubiquitous phenomenon in many guises. Rev. Gen. Psychol. 2, 175–220 (1998). [Google Scholar]
- 109.Rest J. R., Development in Judging Moral Issues (University of Minnesota Press, 1992). [Google Scholar]
- 110.OpenAI, GPT-4 Technical Report. arXiv [Preprint] (2023). 10.48550/arXiv.2303.08774 (Accessed 26 June 2024). [DOI]
- 111.Meta AI, Llama 3: A collection of foundation language models. GitHub. https://github.com/meta-llama/llama3. Accessed 5 Oct 2024.
- 112.Gomila R., Logistic or linear? Estimating causal effects of experimental treatments on binary outcomes using regression analysis J. Exp. Psychol. Gen. 150, 700–709 (2021). [DOI] [PubMed] [Google Scholar]
- 113.H. Singmann, B. Bolker, J. Westfall, F. Aust, M. S. Ben-Shachar, afex: Analysis of Factorial Experiments. (Version 1.1-1, CRAN, 2022).
- 114.Shariff A. F., Does religion increase moral behavior? Curr. Opin. Psychol. 6, 108–113 (2015). [Google Scholar]
- 115.Kivikangas J. M., Fernández-Castilla B., Järvelä S., Ravaja N., Lönnqvist J. E., Moral foundations and political orientation: Systematic review and meta-analysis. Psychol. Bull. 147, 55–94 (2021). [DOI] [PubMed] [Google Scholar]
- 116.A. Alhassan, J. Zhang, V. Schlegel, “‘Am I the bad one’? Predicting the moral judgement of the crowd using pre-trained language models” in Proceedings of the Thirteenth Language Resources and Evaluation Conference (European Language Resources Association, Marseille, France, 2022), pp. 267–276.
- 117.M. Maier, V. Cheung, F. Lieder, Code and data for analyses in “Large language models show amplified cognitive biases in moral decision-making.” Open Science Framework. https://osf.io/3kvjd/. Deposited 4 November 2025. [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Appendix 01 (PDF)
Data Availability Statement
Anonymized data and materials data have been deposited in OSF (https://osf.io/3kvjd/) (117).





