Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2026 Mar 4;123(10):e2519941123. doi: 10.1073/pnas.2519941123

Moral stereotyping in large language models

Aliah Zewail a, Alexandra Figueroa b, Jesse Graham c, Mohammad Atari a,1
PMCID: PMC12974530  PMID: 41779787

Significance

Large language models (LLMs) are increasingly used not only for communication, but also for research tasks like estimating public opinion and simulating moral judgments. But can they truly reflect the diversity of human values across cultures? We compare LLM-generated moral value estimates to real-world survey data from 48 countries and find consistent biases: LLMs overemphasize moral concerns common in Western societies and underestimate values more prominent elsewhere. These distortions likely stem from cultural biases in training data and carry societal implications and risks. Morality is foundational to how people express opinions, justify laws, and engage in politics; thus, distorted moral representations may lead to mischaracterizations of public sentiment. Rather than offering a culturally neutral lens, current models risk reinforcing stereotypes.

Keywords: morality, large language models, culture, AI

Abstract

Can Large Language Models (LLMs) accurately estimate various societies’ moral values? Here, we query the perceptions of LLMs regarding the moral norms of the “average” person from 48 nations and compare them to a large-scale (n=90,802) survey of six moral values (Care, Equality, Proportionality, Loyalty, Authority, and Purity) from those populations. Our findings indicate that LLMs poorly capture the moral diversity around the globe, systematically overestimating some moral values (particularly Care) and underestimating others (especially Purity). Notably, examining various versions of Generative Pre-trained Transformer (GPT) shows that these LLMs may overestimate the overall moral concerns of some Western countries (e.g., the United States, Canada, and Australia) while underestimating those of non-Western countries (e.g., Nigeria, Morocco, and Indonesia). Our work demonstrates that LLMs are inaccurate generators of cross-cultural estimations in the moral domain; in other words, they stereotype the moral values of non-Western populations in predictable ways. Our results highlight the ethical and epistemic risks of relying on LLMs to estimate the endorsement of moral values around the globe.


Traditionally, social scientists have relied on various methods such as ethnographies, surveys, and experiments to gain cultural knowledge about different populations. Recently, however, a new source of knowledge has emerged: large language models (LLMs). Some researchers have proposed that LLMs can aid in social science research (1), even reliably predicting social science experimental results (2). Moreover, some have argued that LLMs may approximate data in place of human participants in certain domains (3). Indeed, one could ask generative AI models such as ChatGPT about different cultures’ norms, values, traditions, and beliefs instead of collecting data in the traditional way. However, it remains unclear to what extent these models can accurately represent various cultural values. Such accounts rely on the assumption that LLMs have accurate information about a wide range of human populations. In this study, we put this assumption to the test and directly examine how accurately LLMs can estimate the moral values of different nations around the globe.

If LLMs have an accurate representation of diverse cultures’ moral values, they should be able to estimate the moral values of the average* person in those cultural contexts. If LLMs’ estimates are far from the ground truth (actual moral values of those cultures), LLMs are showing evidence of inaccurate stereotyping of a population’s moral endorsements. Stereotyping in humans involves several processes, including social learning, cultural norms, motivational factors, and—most relevant to our theorizing—limited exposure to information or experiences (47). LLMs exhibit an analogous, though mechanistically distinct, process as their training regimen draws from purely linguistic co-occurrence data (i.e., patterns of tokens in massive text corpora), supplying them with salient (but imbalanced) information (8). Thus, their “stereotypes” emerge as a latent linguistic bias, whereby the dominating information embedded in the model comes to define what is normative, while all else is linguistically silenced or misrepresented (9, 10). To say that LLMs “stereotype” is simply to note that they absorb and reproduce the statistical contours of their training environment. Much as humans intuit patterns from repeated exposure, these models treat frequent pairings as typical and rare ones as atypical. But because LLMs, unlike humans, have no knowledge beyond text (e.g., no perception), they cannot correct for distortions in their training data (unless external layers are corrected in fine-tuning or reinforcement learning) (11). What emerges in LLMs is hence a purely linguistic category of stereotyping: LLMs’ sense of “what people are typically like” echoes the most common linguistic phrases in their training data.

As a result, LLMs are likely reflective of a dataset where Western, English texts are overweighted, and thus overrepresented in subsequent outputs (12). This is reflective of common stereotype inaccuracy, where populations are overgeneralized (13). When such correlations reflect social or cultural regularities—for example, bias against older women in the workplace (14)—the model inherits and reproduces those patterns. Researchers have noted (9) that AI systems often instill and project the identity-based biases originally formed in the human data on which they are trained. When these associations are applied too rigidly or without context, the model “stereotypes,” producing biased completions or inferences that mirror overgeneralized human beliefs. Recent work (15) further illustrates that LLMs generate stereotypical outputs across domains such as gender, race, religion, and occupation, due to skewed exposure in their training distributions and reliance on co-occurrence statistics.

In the context of morality, these dynamics imply that LLMs may overweight unrepresentative information to produce distorted estimates of moral values across cultures, reflecting not just inaccuracy, but a process similar to stereotyping in humans. We find the pattern of inaccuracy to be particularly interesting, as it is similar to a type of stereotype commonly found in human reasoning–valence inaccuracy (13). Rather than the estimates being randomly inaccurate, these may models systematically inflate the positive moral traits of the data that are overrepresented in the training data [similar to how people often inflate positive qualities in those that are culturally similar to them (16)] and underestimate positive qualities in those that are least represented in their training regimen, following a similar pattern to which humans often enact. Importantly, ad-hoc debiasing methods (10) [e.g., Reinforcement Learning from Human Feedback; RLHF (17)] may aggravate these probability distributions as the models are disproportionately exposed to Western, English-speaking, and elite cultural inputs (8, 18). Furthermore, stereotypes can fuel divisions within social dynamics (19) and international relations. For example, ref. 20 analyzed how American liberals and conservatives stereotype their own ideological group and the opposing side. Comparing judgments generated by individuals about their own morals to those generated by political opposites about “typical” liberals or conservatives, the authors found that both liberals and conservatives exaggerate how different their moral values were from one another. This overestimation bias may stem from an underestimation of shared values between the parties (20).

There are a number of typologies for the study of moral values. One of the most influential is Moral Foundations Theory [MFT; (21, 22)], a pluralistic framework whose latest version offers six basic dimensions of moral judgment: Care, Equality, Proportionality, Loyalty, Authority, and Purity (23). Respectively, these six values encompass the virtues of compassion, egalitarianism, merit, solidarity, deference, and piety, and the vices of cruelty, inequality, disproportionality, betrayal, subversion, and degradation. As globalization increases, these cultural elements are transmitted through more complex interactions occurring within increasingly diverse societies, although that may not necessarily mean that countries will converge on a set of values. Contact between those with differing moral frameworks is sometimes simplified by stereotyping, that is, reducing otherwise complex and diverse cultures into what one might consider “typical” (13, 24). While stereotypes can be a helpful heuristic for interacting with unfamiliar others, they can also be highly inaccurate, and have the potential to come at a social cost. One type of stereotype inaccuracy—valence inaccuracy—can manifest as a relative overestimation of positive traits for one’s in-group (positive valence inaccuracy) and/or an overestimation of negative traits in an outgroup (negative valence inaccuracy) (13, 25). Such a misperception can lead to deeper and more dangerous social cleavages, especially in the moral domain (20).

As LLMs become increasingly integrated into societies, they too, alongside humans, are becoming part of these complex social interactions. Some researchers have argued that they are less biased than human participants and are perhaps better generators of predictive data (3). However, LLMs’ perception of different cultures may be skewed as most of their training data come from a thin slice of human diversity around the globe. These texts predominantly reflect WEIRD [Western, Educated, Industrialized, Rich, and Democratic; (26)] perspectives, as less-WEIRD populations have lower levels of internet output and thus are not a primary data source for AI companies (8).

Although some researchers have likened the cognitive and attitudinal characteristics of LLMs broadly to those of “humans,” research suggests that LLMs have mainly acquired the psychology of WEIRD humans (8). Indeed, LLMs hold geographical biases against poorer nations on a range of subjective (e.g., likeability, attractiveness, and intelligence) topics, notably rating these characteristics much lower than nations of higher socioeconomic status, like European nations (27). We propose that the inherent WEIRD bias in LLMs, likely rooted in their training corpora, may lead to a skewed representation of moral values across all populations. Specifically, we investigate the degree to which different versions of GPT (i.e., GPT-3.5, GPT-4, and GPT-4o) inaccurately generate the moral values of 48 different nations. Researchers have demonstrated that one glaring issue of LLMs is their disproportionate favoritism of cultural values of English-speaking and Protestant European nations and their cultural misalignment with less-WEIRD populations (18). These misrepresentations and biases are being increasingly adopted by researchers and already consumed by millions of laypersons. As LLMs progressively become more involved in some human decision-making processes, examining potential inaccuracies within a model’s output prompts ethical AI development and usage, enhancing the reliability, accuracy, and cultural competence of what is an emerging pillar in society.

The Present Study.

We examine the six moral foundations of 48 nations computed by three different versions of Open AI’s GPT family of LLMs (GPT-3.5, GPT-4, and GPT-4o) to investigate whether these models accurately represent the moral values of various nations. We also replicate our findings using moral value estimates generated by Meta’s LLaMa (versions 2-7b and 2-70b-chat-hf) and Google’s Gemini Pro (SI Appendix).

Methods

Human Participants.

It is important that “synthetic” estimates of LLMs be benchmarked against a ground truth. As an accuracy benchmark, the moral value scores generated by our three versions of GPT were compared to human participant scores sourced from the YourMorals database for each country (median sample size per country = 299, total N=90,802). The Institutional Review Board (IRB) at the University of Southern California approved this study (UP-07-00393). These values were sourced from YourMorals.org, an online platform that provides users with feedback on their moral values, and, in turn, contributes to a large-scale database of anonymized responses. Data are collected using the Moral Foundations Questionnaire-2 (MFQ-2) (23), a psychometrically validated instrument that quantifies endorsement of the six aforementioned moral values. Previous research has drawn on this dataset in various applications, ranging from behavioral interventions related to climate change (28), predicting county-level COVID-19 vaccination rates in the United States. (29), temporal trends in moral values (30), and studying voting behavior in the United States (31), demonstrating its value as a credible tool for measuring moral values.

The primary measure of YourMorals, MFQ-2, is a 36-item self-report measure wherein participants (after obtaining their consent) indicate their agreement with items related to the moral foundations of Care (e.g., “Caring for people who have suffered is an important virtue”), Equality (e.g., “I believe that everyone should be given the same quantity of resources in life”), Proportionality (e.g., “I think people should be rewarded in proportion to what they contribute”), Loyalty (e.g., “Everyone should love their own community”), Authority (e.g., “I think it is important for societies to cherish their traditional values”), and Purity (e.g., “I think the human body should be treated like a temple, housing something sacred within”). Participants respond along a 5-point Likert-type scale, ranging from 1 (Does not describe me at all) to 5 (Describes me extremely well). Notably, all participants completed this measure online and in English, which may breed a bias in which our participants represent more educated, globally connected, and socioeconomically privileged subgroups within their respective nations. To help combat this potential bias from a nonrepresentative pool, we performed Multilevel Regression with Synthetic Poststratification [MrsP; (32)] to ensure more representative national averages (33). Specifically, we fit multilevel models predicting the endorsement of the six moral foundations using individual-level indicators for age, sex, and country. Then, we poststratified the predictions using World Bank census data that report the distribution of age and sex for each country. Incorporating this method yields country-level estimates that better reflect the demographic structure of our 48 countries, thus reducing sampling bias (34). We refer to these data points as the “actual scores” on moral values.

LLM Procedures and Prompts.

Our GPT estimates were sourced from three versions of the language model: GPT-3.5, GPT-4, and GPT-4o. The GPT-3.5 model was a refined iteration of the GPT-3 model. The GPT-4o model was designed to provide GPT-4-level intelligence with enhanced speed and improved capabilities. We filtered the YourMorals data to only include countries for which we had 100 participants or more, leaving us with 48 countries. Our GPT queries were also done with these 48 countries. To provide some variance and serve as a proxy for LLMs’ confidence in their response, we queried GPT 10 times for each question.

For consistency and similarity between our human sample and LLM estimations, we prompted each version of GPT in English with each survey item, adapted from the MFQ-2, one after another making sure to use identical prompting strategies, hyperparameters, and sampling methods (SI Appendix, Table S2). Additionally, we prompted the three iterations of GPT with two variations of the same task. We asked all 36 items in the MFQ-2 for the “average” and “random” person in each country to ensure the reliability of our results (13, 35). All models of GPT ran through 10 iterations of the prompt below using two different prompt variations, resulting in 20 estimations of each item for each country per version of GPT, which were averaged. In total, our GPT dataset contains 103,680 data points (48 countries × 6 foundations × 6 items in each foundation × 2 different prompts × 3 different GPT versions × 10 iterations for each item). Our code is publicly available here: https://osf.io/afq67/?view_only=64f2f6b78b6e4f0abb6e441e567b9ebe.

“For the statement below, please indicate how well the statement describes the [average/random] person from [country]. Response options: Does not describe the [average/random] person at all (1); slightly describes the [average/random] person (2); moderately describes the [average/random] person (3); describes the [average/random] person fairly well (4); and describes the [average/random] person extremely well (5). Please answer only using a single number, with no words.”

We refer to these data points as the “GPT scores.” Because of the low Cronbach’s alphas of GPT foundation scores generated with the “random” prompt (SI Appendix), we exclusively use GPT scores generated using the “average” prompt. The internal consistencies for GPT-3.5, GPT-4, and GPT-4o were as follows (respectively): Care (α=0.72; α=0.85; α=0.88), Equality (α=0.69; α=0.74; α=0.82), Proportionality (α=0.71; α=0.72; α=0.65), Loyalty (α=0.73; α=0.84; α=0.85), Authority (α=0.85; α=0.89; α=0.95), and Purity (α=0.83; α=0.83; α=0.84). Visual comparisons across the different prompts for each moral foundation can be found in SI Appendix.

Results

The Inaccuracy of LLMs.

We first evaluate the claim that LLMs can reliably approximate human moral values. First, we take the difference between the GPT scores produced for each foundation against the actual estimates (see Fig. 1 for a visual representation). If GPT were an accurate generator of cultural data, Fig. 1 would be mostly white, indicating no difference between the scores produced by GPT and actual estimates across countries. However, as can be seen, two key issues emerge, highlighting how GPT systematically stereotypes the morality of nations. First, as newer versions of GPT are released, we observe that Equality and Purity are increasingly underestimated in the majority of populations. Second, GPT-4o shows signs of improved accuracy, but the pattern of inaccuracy in underestimating Equality and Purity remains consistent. In fact, we find that this slight improvement in accuracy is primarily seen in nations that are culturally similar to the United States [i.e., WEIRD populations; (36)]. Contrastingly, nations more culturally distant from the United States, i.e., less-WEIRD populations, continue to show a pattern of increasingly being misrepresented in the moral domain across versions of GPT. For example, in our sample, Egypt was culturally furthest away from the United States. All GPT versions underestimated almost all moral concerns (apart from Authority) of the “average” Egyptian (Fig. 1). The discrepancies illustrated in Fig. 1 highlight a general misalignment of LLM estimates and actual human values across the globe, demonstrating how poorly GPT generates cultural data. The magnitude of these differences is consequential; for example, a 1-point difference on this scale for Equality is about the difference observed between Egypt and New Zealand, also comparable with liberal-conservative differences in the United States (37). Moreover, to further quantify these discrepancies, we computed Cohen’s d with 95% CIs and found large effect sizes (Care (1.01 [0.81, 1.20]), Equality (0.45 [0.65, 0.25]), Proportionality (0.42 [0.23, 0.60]), and Loyalty (0.58 [0.43, 0.73]), Authority (1.00 [0.82, 1.18]), and Purity (0.84 [0.96, 0.71]).), indicating that gaps between scores generated by LLMs and humans translate to substantively meaningful distortions. These differences are comparable to or greater than between-country variations reported in previous research with MFT (23, 38, 39). Results using the moral value estimates computed by Meta’s LLaMa and Google’s Gemini Pro are present in SI Appendix.

Fig. 1.

Comparison of G P T (3.5, 4, 4o) moral estimates vs. human ground truth. Red indicates overestimation; blue indicates underestimation.

The difference between the moral value estimate (produced by GPT-3.5, GPT-4, and GPT-4o) and the ground truth (human sample). Red coloration represents an overestimation of moral values, and blue coloration represents an underestimation of moral values.

Cross-National Discrepancies.

The magnitude of moral stereotyping (i.e., discrepancies between GPT and actual scores) was quantified by distance scores—measured by the Euclidean distance between GPT and actual scores in a six-dimensional space—to capture the overall inaccuracy of GPT moral estimates, and difference scores—measured by the difference between GPT and human moral estimates—to assess the direction of LLMs’ stereotyping (overestimation vs. underestimation). The Euclidean distance was calculated for each nation to quantify the discrepancy between all six moral values estimated by LLMs and endorsed by humans in our cross-national data (Formula 1). In this formula, q represents the set of moral values computed by the different models of GPT, p denotes the corresponding human-reported benchmark, and i represents each of the six moral foundations. For each nation (j), we summed the squared differences between GPT and human scores for all moral foundations (i), then computed the square root to obtain the Euclidean distance (d). To conceptually express what Euclidean distance captures, we refer to it as a measurement of models’ inaccuracy of the overall moral concern of a nation. Results for the latest version of GPT are shown in Fig. 2, highlighting GPT-4o’s substantial deviations from the actual moral values across most countries (t=35.61,df=47,MeanEuclideanDistance=1.43, 95%CI [1.35, 1.51], P<0.001), demonstrating a general misalignment with global moral norms that is particularly prominent in Middle Eastern and Sub-Saharan African nations. Moreover, as we can see in Fig. 2, GPT-4o was more accurate in estimating the overall moral concerns of WEIRDer nations, that is, countries that are culturally closer to the United States as the prototypical WEIRD nation (36), r= 0.56, P< 0.001.

dj=i=16(qijpij)2 [1]

Note. This equation computes the inaccuracy of overall moral concern as the Euclidean distance between two entities in a six-dimensional moral value space. In our work, we used it to calculate the distance between the moral value estimates produced by each GPT model and the corresponding estimates derived from human participants.

Fig. 2.

Thematic map of the world with a score from one point zero zero to two point two five.

The map above illustrates, overall, how much GPT-4o’s moral endorsements deviate (i.e., distance score) from the actual participant scores for each country.

Underestimation vs. Overestimation.

Stereotype valence inaccuracy, as previously described, can result in either an overestimation or underestimation of various characteristics, with positive characteristics usually being overestimated in one’s own group and underestimated in outgroups (13, 25). Given that general moral concern is typically considered a positive trait (40, 41), we considered that if LLMs are biased, they may overestimate moral concern in some cultural groups while underestimating it in others. To examine the valence of these inaccuracies by moral foundation and GPT version, we calculated a series of correlations between the WEIRD distance scores [cultural and psychological distance from the United States; (36)], and compared them to the difference between the GPT and actual scores for 37 of the 48 countries (11 countries did not have WEIRD distance scores). Results are shown in Fig. 3. We find evidence that while there is some variation in different versions of GPT, these models are systematically biased, often overestimating moral values in countries such as the United States, Canada, and Australia, while underestimating the moral concerns of the Middle East and Africa.

Fig. 3.

Plots of G P T vs. human moral scores. G P T overestimates Care, Loyalty, and Authority; underestimates Equality and Purity.

The plots above visually break down the correlations between difference scores and cultural distance scores by the GPT model. Also, GPT generally tends to overestimate four values (Care, Proportionality, Loyalty, and Authority) while underestimating two values (Equality and Purity).

Our results partially confirm a WEIRD bias in GPT, as at least one version of this LLM inaccurately enhances the perceived morality of WEIRDer populations, portraying them as more morally concerned than what is reflected by their actual scores along some moral dimensions, namely Care, Equality, and Proportionality. However, across all moral foundations and all GPT models, there is a consistent pattern of relative over/underestimation, such that as a nation becomes more culturally similar to the United States, GPT rates its moral concern on each foundation as higher than those nations that are less culturally similar to the United States (excluding Purity scores, which are notably less normative in western cultures (23, 42).

Determinants of Moral Stereotyping in LLMs.

Although we previously provided evidence pointing to a WEIRD bias, more should be done to unpack the kinds of non-WEIRD nations that get stereotyped by LLMs. As Henrich (43) have argued, psychological variation does not fall along some linear continuum from WEIRD to non-WEIRD. For example, Norway and Nigeria are almost equally distant from the United States but are obviously quite different in terms of their culture, history, and economy. What kinds of non-WEIRD countries are more prone to be morally stereotyped by LLMs? As an exploratory analysis, we investigated different country-level measures of industrialization relevant to data availability to explain the overall magnitude of moral stereotyping. Specifically, we examined Gross Domestic Product at Purchasing Power Parity, GDP (PPP) per capita, press freedom, and population as factors influencing LLMs’ moral values output. We theorize these factors may influence what and how much data is available for AI companies to use to train their models. First, richer countries have higher internet usage and contribute more content to the internet. Thus, a higher GDP should be associated with a larger internet footprint, therefore providing AI companies with more extensive data to train their models, potentially resulting in more precise moral value estimates for that nation. Second, greater press freedom may result in higher levels of information access on the internet. To capture this factor, we standardized the Press Freedom Index from Reporters Without Borders which measures the level of freedom available to journalists in various countries (0 indicates the lowest press freedom, and 100 indicates the highest press freedom). Third, the larger a country’s population, the more individuals it has to generate online content. The number of people residing in a nation may influence the accuracy of these moral value estimates as the more people there are, the more data present for these models to train on [although research suggests there is a misalignment between the world’s most spoken languages and those most represented in the GPT training corpus; (44)]. For GPT-3.5, zero-order correlations suggested a surprising positive correlation between economic development and moral stereotype inaccuracy (r=0.40,P=0.005), while there was no significant correlation between press freedom (r=0.23,P=0.117) and population (r=0.13,P=0.361), with moral stereotype inaccuracy. However, after putting all these variables in an Ordinary Least Squares (OLS) regression and accounting for the nonindependence of nations by adding a fixed effect for the continent, none of the predictors significantly explained GPT-3.5’s moral stereotype inaccuracy (Table 1). For GPT-4, the magnitude of inaccuracy was smaller for countries with greater economic development (r=0.30,P=0.036) and press freedom (r=0.54,P<0.001) and larger for more populous countries (r=0.41,P=0.003). In the OLS regression, only free press was found to be significantly associated with moral stereotype inaccuracy (Table 1), such that countries with greater press freedom were stereotyped to a lesser degree by GPT-4. The GPT-4o model showed a similar pattern to GPT-4: Stereotype inaccuracy was smaller for countries with greater economic development (r=0.40,P=0.005) and press freedom (r=0.64,P<0.001), but the effect of population size was not significant (r=0.26,P=0.069). As can be seen in Table 1, only press freedom was a significant predictor of GPT-4o’s moral stereotype inaccuracy in a regression (B=0.14,SE=0.047,P=0.005). Of note, the multidimensional Euclidean metric for moral stereotype inaccuracy quantifies divergence from reality, but does not indicate the directionality of this inaccuracy (for directions, see Fig. 3).

Table 1.

Three regression models to predict the magnitude of inaccuracy in estimating moral values for three LLMs

Inaccuracy of overall moral concern
GPT-3.5 GPT-4 GPT-4o
GDP (PPP) per capita 0.105
(0.155)
0.140
(0.097)
0.120
(0.073)
Press freedom 0.098
(0.101)
−0.140∗
(0.063)
−0.142∗∗
(0.047)
Population 0.114
(0.175)
0.176
(0.109)
−0.005
(0.082)
Continent FE Included Included Included
Constant 1.253∗∗∗
(0.190)
1.519∗∗∗
(0.119)
1.901∗∗∗
(0.089)
Observations 48 48 48
R2 0.272 0.649 0.673
Adjusted R2 0.122 0.577 0.606
Residual SE (df = 39) 0.374 0.233 0.175
F Statistic (df = 8; 39) 1.820 9.025∗∗∗ 10.033∗∗∗

A fixed effect (FE) was added to account for the clustering of nations in continents.

Note: P < 0.05; ∗∗ P < 0.01; and ∗∗∗ P < 0.001.

Robustness Checks.

To ensure our results are not artifacts of the linguistic structure or theoretical framing of our primary analyses with MFQ-2, we conducted two independent replication studies. Importantly, each replication addresses distinct sources of potential bias in our main analyses. The first examines whether our results persist when human moral values are measured in a population’s local language, addressing concerns about linguistic or sampling biases in the English-only YourMorals database. The second investigates whether the observed moral discrepancies between LLMs and humans uphold under a theoretically distinct model of morality, ensuring that our findings are not tied to the structure of Moral Foundations Theory alone. Taken together, these two studies draw on new data, new languages, and a new theoretical framework, providing rigorous cross-validation to demonstrate the stability of our findings.

Cross-linguistic replication using translated moral foundations data.

A potential limitation of our primary analyses is that all MFQ-2 data were collected in English across 48 nations, which may systematically bias our “ground truth” benchmarks because those who complete English-language surveys may represent a more educated, globally connected, or socioeconomically privileged subgroup, rather than the broader population. To address this concern, we conducted a large-scale cross-cultural replication designed to eliminate the linguistic channel as a possible confound.

We administered a validated short form of the MFQ-2 (MFQ-2-S; 45), translated into six languages (Arabic, Urdu, Ukrainian, Portuguese, Spanish, and Turkish). Translations were reviewed by language experts, ensuring that conceptual meaning of each survey item was preserved.

We collected 4,666 human responses across 11 populations, targeting a more demographically diverse pool than online English-language samples. Countries were selected specifically to capture variance in WEIRDness, economic development, political structure, and cultural distance from the United States: Argentina (n=119), Brazil (n=327), Colombia (n=473), Egypt (n=302), India (n=606), Mexico (n=150), Pakistan (n=483), Turkey (n=635), Ukraine (n=320), the United Kingdom (n=716), and the United States (n=855). In parallel, we prompted GPT-3.5, GPT-4, and GPT-4o to estimate the moral endorsements of these countries in their respective languages, removing the possibility that English-language prompting alone shaped the models’ responses.

The detailed findings of this additional study are available in SI Appendix. In short, the results remain robust. Across translated versions of the MFQ-2-S, GPT models continued to systematically underestimate endorsements of Equality and Purity, particularly in less-WEIRD populations (e.g., Egypt and Turkey). Moreover, in the world map where GPT4o’s overall inaccuracy of moral concern is shown (SI Appendix, Fig. S8), the same cross-cultural asymmetry persists: Moral values in WEIRDer nations are estimated more accurately compared to populations in less-WEIRD contexts. Additionally, with this independent dataset, GPT models inflated the moral values of WEIRDer countries (compared to less-WEIRD populations), reproducing the same directional bias we observed in our original analyses. Last, although the coefficient for press freedom is no longer statistically significant in this smaller sample, its magnitude and directionality remain consistent (refer to SI Appendix, Table S10). This reaffirms that nations with higher levels of press freedom—a characteristic of WEIRD societies—display higher accuracy for the overall moral concern of that population.

This replication demonstrates that our findings are not artifacts of English-language or MFQ-2 item-wording. When we remove language constraints by collecting new data in native languages and prompting LLMs equivalently, the same moral discrepancies in LLM accuracy emerge, suggesting that the source of their misestimation lies not in the linguistic structure of MFQ-2, but in the cultural structure of the models themselves.

Testing generalizability across moral frameworks.

A second concern is that our primary results may hinge on the specific moral taxonomy of Moral Foundations Theory, which, given its well-established centrality in moral psychology (21, 22), served as the basis for our main analyses. Crucially, if our findings reflect stereotyping within the theoretical categories embedded in MFT, they may not generalize to other organizations of morality. Thus, to test whether our conclusions remain beyond our current framework, we replicated our study using the Morality-as-Cooperation (MAC) framework (46), which conceptualizes morality in terms of seven cooperative strategies: family, group, reciprocity, bravery, respect, fairness, and property (46, 47). Unlike MFT, which groups moral concerns into intuitive foundations, MAC forms its structure from evolutionary game theory, providing a conceptually distinct lens for evaluating the moral representations embedded in LLMs.

We leveraged the nationally representative dataset from ref. 48 which includes moral endorsement scores of 63 countries (29 different languages) measured using the MAC questionnaire (MACQ; 49). To compare these data, we prompted GPT 3.5, GPT-4, and GPT-4o to complete the same MACQ in the same countries and languages.

Analyses (reported in SI Appendix, Figs. S9 and S10) reveal similar patterns of moral stereotyping in less-WEIRD nations. Specifically, all three versions of GPT provide moral estimations that are most accurate for nations culturally proximate to the United States and they deviate substantially for less-WEIRD populations, mirroring our main analyses.

This conceptual replication provides robustness to the patterns we report above, highlighting that our results are not confined to Moral Foundations Theory or the structure of MFQ-2. Rather, it further reflects the cultural representations embedded within these AI systems.

Taken together, our linguistic and theoretical replications provide a robust cross-validation of our findings: LLMs systematically provide inaccurate moral estimates for diverse populations around the globe, persisting across languages and theories.

Discussion

As LLMs continue to expand their purpose globally (50), it is crucial to evaluate the quality of their knowledge and their effects on scientific discovery (51, 52). This is particularly pressing given that some researchers have argued that AI will transform social science research (1, 53), and some called for the use of LLMs as data generators, potentially replacing human participants (3, 54). In the realm of morality, Dillion et al. (3) have argued that not only can we replace humans with LLMs as participants, but perhaps we ought to, crowning LLMs as the paragons of ethics (see ref. 55). However, a simple yet fundamental issue remains: Which humans are (over-)represented in LLMs? And whose moral compass are these models calibrated to? Here, we demonstrate a general inaccuracy in GPT’s estimations of moral values across 48 countries. The inaccuracies are systematic and can be best described as a form of stereotyping: valence inaccuracy.

Past research has demonstrated that LLMs are aligned with a WEIRD psychology (8), leading this purportedly “global” tool to often misrepresent “humans” or stereotype many “people” around the world. These authors argued that there is a WEIRD-in-WEIRD-out (WIWO) bias in these models: Since less-WEIRD viewpoints, languages, and demographics are underrepresented in these models’ training data, their outputs may marginalize or misrepresent those perspectives. In other words, the imbalance within their training data serves as a function of limited exposure. Stereotypes are associations between social groups and semantic attributes that are widely shared within societies, and arise in written language corpora (56, 57). Just as human stereotypes arise when individuals rely on salient (but unrepresentative) information, LLMs are likely to overgeneralize from the disproportionately WEIRD content in their training data (8). Here, we demonstrated that LLMs are partially inaccurate in their estimations of less-WEIRD countries across some moral dimensions, particularly Care, Equality, and Proportionality. Additionally, GPT models were most likely to underestimate endorsements of Equality and Purity across a majority of our 48 nations. Importantly, we believe that postulating a coherent explanation of the observed patterns requires clearer connections between the moral composition of LLMs training corpora and their behavior. For instance, if the discrepancies were limited to underestimations in Purity (the least WEIRD-aligned moral foundation), one might attribute them to a WEIRD-skewing of LLMs’ training data. However, this does not readily account for why Equality, a value commonly emphasized in Western contexts, would be underestimated as well. For this reason, we view the observed pattern as an intriguing and promising venture for future research in cultural and moral psychology and as a clear motivation for greater transparency in AI development.

When LLMs carry inaccurate moral priors into everyday tools, the consequences show up in quiet but important ways, influencing how public opinions and attitudes are represented worldwide. For instance, plausible downstream implications may include a mental-health AI assistant used in East Asian contexts might prioritize individual autonomy over filial obligations, steering family conflict advice toward “setting boundaries” rather than negotiated harmony (related to the foundation of Loyalty). Moreover, automated hiring screeners that rate narratives of “impact” and self-promotion more highly can disadvantage candidates from cultures that value tradition and collective credit, even when qualifications match (related to the foundation of Authority). Another possible example can be found within content moderation systems trained on Western norms flagging depictions of animal slaughter during Eid al-Adha or open-air cremation rituals as harmful or obscene, while permitting equally graphic but culturally familiar content, which tilts online visibility (related to the foundation of Purity). Additionally, public-health chatbots that cast vaccination solely as personal risk management can miss community duty narratives that motivate uptake in high-purity regions (29). Not all these failures are catastrophic on their own, but together they homogenize global moral diversity, create small systematic disadvantages for some groups, and might lower trust in AI systems across cultures (58).

Additionally, research showcases that LLMs rely on shared latent representations across an array of tasks, such that biases in their internal value structure may not be isolated to the task of MFQ-2 or MACQ outputs, but are recycled across different reasoning, expression, and summarizations (59). For example, recent research indicates that LLMs may structure moral judgments within their embedding spaces in ways that parallel human perceptions of right and wrong (60). As such, the cultural orientation of LLMs (which favors the values of English-speaking Protestant European countries) can shape user expressions (61) and short-term attitudes (62), which may accumulate over time within cultural systems (63). Thus, the misrepresentation of moral estimates in LLMs may pose risks in downstream tasks that involve right and wrong: As these models are increasingly deployed in tools that summarize public attitudes, support decision-making, or simulate stakeholder perspectives, systematic distortions in moral representations could contribute to a skewed portrayal of entire populations, misinforming users and policymakers alike (64). The practical remedy may lie in expanding the training corpora, culturally stratified evaluation sets, regional fine-tuning, and explicit checks for framing parity before deployment.

To further investigate the downstream implications of how moral and broader cultural representations of populations encoded in LLMs translate into real-world applications, we encourage future research to directly compare model-generated population inferences on policy-relevant moralized issues against human benchmarks. Ref. 65 offers a blueprint to this approach, assessing how LLMs fail to replicate American moral judgments using concrete moral scenarios. This work is especially timely because prior work shows that fine-tuning and alignment interventions (ranging from bias-inducing fine-tuning and RLHF to direct manipulation of internal representations) can systematically alter LLMs’ expressed political leanings and downstream behavior, demonstrating that such normative attributes are malleable rather than fixed (6668).

We note that Moral Foundations Theory is one typology among many within the field and that our use of MFQ-2 was anchored on the fact that it is a validated psychometric scale that has been used in numerous studies (29, 69, 70). Thus, we encourage future work to replicate our results using other theories on morality, such as the Relationship Regulation Theory (71), Schwartz’s Values Model (72), Model of Moral Motives (73), Theory of Dyadic Morality (74) or the Morality-as-Cooperation (47). Additionally, although the development of MFQ-2 included validation with non-WEIRD populations (23), it remains possible that certain item wordings may not work as well in some cultural contexts (75, 76). However, our findings, based on translated versions of the MFQ-2, suggest that such bias is unlikely to account for the observed patterns. We argue that these effects are likely to replicate across moral typologies, and may have even further applications to other traits, as these results are likely driven by LLMs’ (lack of) access to cultural data from less-WEIRD populations.

As an exploratory analysis, we also examined what country-level factors might contribute to the discrepancy between GPT’s estimations and actual moral values worldwide. Overall, GPT’s accuracy was shown to be higher in countries with high levels of press freedom—a characteristic of WEIRDer nations. This aligns with our finding that GPT exhibits some WEIRD-leaning tendencies and is consistent with prior research (8, 18, 44). These findings make clear that GPT is more likely to suppress the moral concerns of less-WEIRD populations while inflating its estimates of WEIRDer populations. Such societies are already underrepresented in social science research (77), and researchers aiming to accurately estimate the moral values of “humans” should not rely on the purported accuracy of LLMs.

Moral representations are fundamental to how people perceive social groups and whether these groups engage in cooperation or conflict (78). As illustrated in ref. 20, systematic distortions in perceived moral differences between ideological groups contribute to polarization. This could, in turn, lead to stereotyping and miscommunication between groups. Our results depict an analogous process in LLMs: Models systematically misrepresent the moral concerns of less-WEIRD populations. This work provides empirical evidence that moral predictions from generative AI systems are not uniformly accurate across values or populations, but structured by the moral distributions likely reflected in their training data. Recognizing this pattern is critical for two reasons: a) it clarifies what kinds of inferences researchers can and cannot draw when using LLMs as cognitive models (79) or “synthetic” simulations of human moral judgments (80), and b) it identifies a measurable pathway by which cultural bias may enter value-sensitive technologies. In short, the current work shows that moral estimation (accessible information) in LLMs is neither random nor universal but culturally patterned, a finding that directly informs both the scientific use of LLMs and the design of more accurate AI systems.

In a nutshell, we recommend studying more diverse samples of living humans around the world to understand “human” psychology rather than attempting to distill human understanding from biased off-the-shelf models like GPT. In addition to the raw language model (i.e., the way it is trained and the corpora it uses), how its guardrails are implemented (e.g., through RLFH) may also influence its understanding of not just a society’s average moral values, but cultural information in general (81). In many commercial LLM pipelines, the human annotators who partake in feedback during RLHF are often pooled from Western liberal-democratic contexts. Notably, many benchmarks for AI alignment consist of datasets around harm, created by Western companies and research teams (e.g., ref. 82). As a result, the model may favor responses that align with those normative viewpoints (83). This can create a second filter on top of the training data; even if a model encountered texts from less-WEIRD contexts in pretraining (hence encountering a more diverse assemblage of moral values), the RLHF step may dampen or override outputs that deviate from the dominant value-system (84), further amplifying the WEIRD norms embedded within LLMs training data (10).

Furthermore, past research has shown (and our results echo) that LLM-estimated measures are not reliable and can lead to a lack of knowledge or even create an illusion of understanding and advancement in science (52). One effective remedy for addressing this issue could be diversifying the training data for LLMs by incorporating a broader range of language content from different world regions and different time eras. Equally important is the need for greater transparency and open documentation from AI developers (85). Without access to the composition of training data, it remains unclear whether these observed biases (e.g., underestimations of endorsements of Purity and Equality in less-WEIRD nations), stem from underrepresentation or a dominance of particular ideological framings. Publishing key characteristics of a dataset (e.g., its origin, makeup, the demographics of the individuals it pools from, and descriptions of the processing steps followed to curate it) provides the means necessary to assist researchers in understanding the conduct of AI systems (86). Within the scope of the current work, this level of transparency would allow researchers to trace the conception of these moral distortions, measure the distribution of moral discourse across cultures embedded within LLMs, and evaluate whether specific domains (such as social platforms or news media) are overrepresented in shaping the model’s internal value structure. Most importantly, this level of transparency would equip researchers to progress beyond documenting inaccuracies to diagnosing their respective mechanics and designing culturally inclusive remedies.

Supplementary Material

Appendix 01 (PDF)

Acknowledgments

This material is based upon work supported by the NSF Graduate Research Fellowship awarded to Aliah Zewail under Grant No. 2439846. Any opinion, findings, and conclusions or recommendations expressed in this material are those of the authors(s) and do not necessarily reflect the views of the NSF. Moreover, we thank Satvika Reddy for her assistance in this project.

Author contributions

A.Z., J.G., and M.A. designed research; A.Z., A.F., and M.A. performed research; A.Z. analyzed data; M.A. supervision; and A.Z., A.F., J.G., and M.A. wrote the paper.

Competing interests

The authors declare no competing interest.

Footnotes

This article is a PNAS Direct Submission.

*By “average,” we mean the statistical mean.

Data, Materials, and Software Availability

Code data have been deposited in OSF (https://osf.io/afq67/?view_only=64f2f6b78b6e4f0abb6e441e567b9ebe) (87). All other data are included in the manuscript and/or SI Appendix.

Supporting Information

References

  • 1.Grossmann I., et al. , AI and the transformation of social science research. Science 380, 1108–1109 (2023). [DOI] [PubMed] [Google Scholar]
  • 2.L. Hewitt, A. Ashokkumar, I. Ghezae, R. Willer, Predicting results of social science experiments using large language models. DocSend [Preprint] (2024). https://docsend.com/view/ity6yf2dansesucf. Accessed 20 November 2025.
  • 3.Dillion D., Mondal D., Tandon N., Gray K., AI language model rivals expert ethicist in perceived moral expertise. Sci. Rep. 15, 4084 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Sherman S. J., Sherman J. W., Percy E. J., Soderberg C. K., Stereotype Development and Formation (Oxford University Press, 2013). [Google Scholar]
  • 5.C. S. Crandall, A. Eshleman, L. O’Brien, Social norms and the expression and suppression of prejudice: The struggle for internalization. J. Pers. Soc. Psychol. 82, 359–378 (2002). [PubMed]
  • 6.J. T. Jost, M. R. Banaji, The role of stereotyping in system-justification and the production of false consciousness. Br. J. Soc. Psychol. 33, 1–27 (1994), https://bpspsychub.onlinelibrary.wiley.com/doi/pdf/10.1111/j.2044-8309.1994.tb01008.x.
  • 7.Sherman J. W., Development and mental representation of stereotypes. J. Pers. Soc. Psychol. 70, 1126–1141 (1996). [DOI] [PubMed] [Google Scholar]
  • 8.M. Atari, M. J. Xue, P. S. Park, D. Blasi, J. Henrich, Which humans? PsyArXiv [Preprint] (2023). 10.31234/osf.io/5b26t (Accessed 1 September 2025). [DOI]
  • 9.Kang O., Hirschi K., Bias and stereotyping: Human and artificial intelligence (AI). Annu. Rev. Appl. Linguist. 45, 69–84 (2025). [Google Scholar]
  • 10.Lin Z., et al. , Towards trustworthy LLMs: A review on debiasing and dehallucinating in large language models. Artif. Intell. Rev. 57, 243 (2024). [Google Scholar]
  • 11.Resnik P., Large language models are biased because they are large language models. Comput. Linguist. 51, 1–21 (2025). [Google Scholar]
  • 12.E. M. Bender, T. Gebru, A. McMillan-Major, S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (Association for Computing Machinery, 2021), pp. 610–623.
  • 13.Judd C. M., Park B., Definition and assessment of accuracy in social stereotypes. Psychol. Rev. 100, 109 (1993). [DOI] [PubMed] [Google Scholar]
  • 14.Guilbeault D., Delecourt S., Desikan B. S., Age and gender distortion in online media and large language models. Nature 646, 1–9 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Z. Wu, S. Bulathwela, M. Perez-Ortiz, A. S. Koshiyama, Stereotype detection in LLMs: A multiclass. explainable, and benchmark-driven approach. arXiv [Preprint] (2024). http://arxiv.org/abs/2404.01768arXiv:2404.01768 (Accessed 20 November 2025).
  • 16.M. Hewstone, M. Rubin, H. Willis, Intergroup bias. Annu. Rev. Psychol. 53, 575–604 (2002). [DOI] [PubMed]
  • 17.Y. Bai et al., Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv [Preprint] (2022). http://arxiv.org/abs/2204.05862 (Accessed 20 November 2025).
  • 18.Tao Y., Viberg O., Baker R. S., Kizilcec R. F., Cultural bias and cultural alignment of large language models. PNAS Nexus 3, pgae346 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.A. J. Cuddy, S. T. Fiske, P. Glick, “Warmth and competence as universal dimensions of social perception: The stereotype content model and the BIAS map” in Advances in Experimental Social Psychology, M. P. Zanna, Ed. (Elsevier Academic Press, 2008), vol. 40, pp. 61–149.
  • 20.Graham J., Nosek B. A., Haidt J., The moral stereotypes of liberals and conservatives: Exaggeration of differences across the political spectrum. PloS One 7, e50092 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.J. Graham et al. , “Moral foundations theory: The pragmatic validity of moral pluralism” in Advances in Experimental Social Psychology, P. Devine, A. Plant, Eds. (Elsevier, 2013), pp. 55–130.
  • 22.Haidt J., Joseph C., Intuitive ethics: How innately prepared intuitions generate culturally variable virtues. Daedalus 133, 55–66 (2004). [Google Scholar]
  • 23.Atari M., et al. , Morality beyond the weird: How the nomological network of morality varies across cultures. J. Pers. Soc. Psychol. 125, 1157–1188 (2023). [DOI] [PubMed] [Google Scholar]
  • 24.Katz D., Braly K. W., Verbal stereotypes and racial prejudice. J. Abnorm. Soc. Psychol. 28, 280–290 (1933). [Google Scholar]
  • 25.Judd C. M., Ryan C. S., Park B., Accuracy in the judgment of in-group and out-group variability. J. Pers. Soc. Psychol. 61, 366 (1991). [DOI] [PubMed] [Google Scholar]
  • 26.Henrich J., Heine S. J., Norenzayan A., The weirdest people in the world? Behav. Brain Sci. 33, 61–83 (2010). [DOI] [PubMed] [Google Scholar]
  • 27.R. Manvi, S. Khanna, M. Burke, D. Lobell, S. Ermon, “Large language models are geographically biased” in Proceedings of the 41st International Conference on Machine Learning, ICML’24 (JMLR.org, 2024).
  • 28.Sinclair A. H., et al. , Behavioral interventions motivate action to address climate change. Proc. Natl. Acad. Sci. U.S.A. 122, e2426768122 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Reimer N. K., et al. , Moral values predict county-level Covid-19 vaccination rates in the United States. Am. Psychol. 77, 743 (2022). [DOI] [PubMed] [Google Scholar]
  • 30.Hohm I., O’Shea B. A., Schaller M., Do moral values change with the seasons? Proc. Natl. Acad. Sci. U.S.A. 121, e2313428121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Enke B., Moral values and voting. J. Polit. Econ. 128, 3679–3729 (2020). [Google Scholar]
  • 32.Leemann L., Wasserfallen F., Extending the use and prediction precision of subnational public opinion estimation. Am. J. Polit. Sci. 61, 1003–1022 (2017). [Google Scholar]
  • 33.Hoover J., Dehghani M., The big, the bad, and the ugly: Geographic estimation with flawed psychological data. Psychol. Methods 25, 412 (2020). [DOI] [PubMed] [Google Scholar]
  • 34.Ebert T., Götz F. M., Mewes L., Rentfrow P. J., Spatial analysis for psychologists: How to use individual-level data for research at the geographically aggregated level. Psychol. Methods 28, 1100 (2023). [DOI] [PubMed] [Google Scholar]
  • 35.A. Figueroa, The verifiability principle: A framework for studying judgment accuracy, accurately. OSF [Preprint] (2024). 10.31219/osf.io/chs9n (Accessed 20 November 2025). [DOI]
  • 36.Muthukrishna M., et al. , Beyond western, educated, industrial, rich, and democratic (weird) psychology: Measuring and mapping scales of cultural and psychological distance. Psychol. Sci. 31, 678–701 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Graham J., Haidt J., Nosek B. A., Liberals and conservatives rely on different sets of moral foundations. J. Pers. Soc. Psychol. 96, 1029 (2009). [DOI] [PubMed] [Google Scholar]
  • 38.Graham J., et al. , Mapping the moral domain. J. Pers. Soc. Psychol. 101, 366–385 (2011a). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Kivikangas J. M., Fernández-Castilla B., Järvelä S., Ravaja N., Lönnqvist J. E., Moral foundations and political orientation: Systematic review and meta-analysis. Psychol. Bull. 147, 55–94 (2021). [DOI] [PubMed] [Google Scholar]
  • 40.Brambilla M., Sacchi S., Rusconi P., Cherubini P., Yzerbyt V. Y., You want to give a good impression? Be honest! moral traits dominate group impression formation Br. J. Soc. Psychol. 51, 149–166 (2012). [DOI] [PubMed] [Google Scholar]
  • 41.Leach C. W., Ellemers N., Barreto M., Group virtue: The importance of morality (vs. competence and sociability) in the positive evaluation of in-groups. J. Pers. Soc. Psychol. 93, 234 (2007). [DOI] [PubMed] [Google Scholar]
  • 42.Graham J., et al. , Mapping the moral domain. J. Pers. Soc. Psychol. 101, 366 (2011b). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.J. Henrich, WEIRD, M. C. Frank, A. Majid, Eds. (MIT Press, 2024), https://oecs.mit.edu/pub/spow8trw.
  • 44.R. L. Johnson et al., The ghost in the machine has an American accent: Value conflict in GPT-3. arXiv [Preprint] (2022). http://arxiv.org/abs/2203.07785 (Accessed 1 December 2024).
  • 45.A. Hajian et al., Measuring moral pluralism efficiently: Developing the short version of the moral foundations questionnaire-2 (MFQ-2-s). Unpubl. manuscript (2026).
  • 46.Curry O. S., Mullins D. A., Whitehouse H., Is it good to cooperate? Testing the theory of morality-as-cooperation in 60 societies. Curr. Anthropol. 60, 47–69 (2019). [Google Scholar]
  • 47.O. S. Curry, Morality as Cooperation: A Problem-Centred Approach, T. K. Shackelford, R. D. Hansen, Eds. (Springer International Publishing, Cham, 2016), pp. 27–51.
  • 48.Azevedo F., et al. , Social and moral psychology of covid-19 across 69 countries. Sci. Data 10, 272 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Curry O. S., Chesters M. J., Van Lissa C. J., Mapping morality with a compass: Testing the theory of ‘morality-as-cooperation’ with a new questionnaire. J. Res. Pers. 78, 106–124 (2019). [Google Scholar]
  • 50.Capraro V., et al. , The impact of generative artificial intelligence on socioeconomic inequalities and policy making. PNAS Nexus 3, pgae191 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Burton J. W., et al. , How large language models can reshape collective intelligence. Nat. Hum. Behav. 8, 1643–1655 (2024). [DOI] [PubMed] [Google Scholar]
  • 52.Messeri L., Crockett M. J., Artificial intelligence and illusions of understanding in scientific research. Nature 627, 49–58 (2024). [DOI] [PubMed] [Google Scholar]
  • 53.Bail C. A., Can generative ai improve social science? Proc. Natl. Acad. Sci. U.S.A. 121, e2314021121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Lehr S. A., Caliskan A., Liyanage S., Banaji M. R., Chatgpt as research scientist: Probing GPT’s capabilities as a research librarian, research ethicist, data generator, and data predictor. Proc. Natl. Acad. Sci. U.S.A. 121, e2404328121 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.M. Crockett, L. Messeri, Should large language models replace human participants? OSF [Preprint] (2023). 10.31234/osf.io/4zdx9 (Accessed 1 December 2024). [DOI]
  • 56.T. E. S. Charlesworth, V. Yang, T. C. Mann, B. Kurdi, M. R. Banaji, Gender stereotypes in natural language: word embeddings show robust consistency across child and adult language corpora of more than 65 million words. Psychol. Sci. 32, 218–240 (2021). [DOI] [PubMed]
  • 57.Nicolas G., Caliskan A., Directionality and representativeness are differentiable components of stereotypes in large language models. PNAS Nexus 3, pgae493 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.N. G. Sadr, S. Heidariasl, K. Megerdoomian, L. Seyyed-Kalantari, A. Emami, “We politely insist: Your LLM must learn the persian art of taarof” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, V. Peng, Eds. (Association for Computational Linguistics, 2025), pp. 1819–1838.
  • 59.J. Wei et al., Emergent abilities of large language models. arXiv [Preprint] (2022). http://arxiv.org/abs/2206.07682 (Accessed 20 November 2025).
  • 60.Schramowski P., Turan C., Andersen N., Rothkopf C. A., Kersting K., Large pre-trained language models contain human-like biases of what is right and wrong to do. Nat. Mach. Intell. 4, 258–268 (2022). [Google Scholar]
  • 61.Goldberg A., Srivastava S. B., Manian V. G., Monroe W., Potts C., Fitting in or standing out? The tradeoffs of structural and cultural embeddedness. Am. Sociol. Rev. 81, 1190–1222 (2016). [Google Scholar]
  • 62.M. Jakesch, A. Bhat, D. Buschek, L. Zalmanson, M. Naaman, “Co-writing with opinionated language models affects users’ views” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Association for Computing Machinery, 2023), pp. 1–15.
  • 63.Thompson B., Kirby S., Smith K., Culture shapes the evolution of cognition. Proc. Natl. Acad. Sci. U.S.A. 113, 4530–4535 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Summerfield C., et al. , The impact of advanced ai systems on democracy. Nat. Hum. Behav. 9, 1–11 (2025). [DOI] [PubMed] [Google Scholar]
  • 65.Grizzard M., et al. , ChatGPT does not replicate human moral judgments: The importance of examining metrics beyond correlation to assess agreement. Sci. Rep. 15, 40965 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.K. Chen, Z. He, J. Yan, T. Shi, K. Lerman, “How susceptible are large language models to ideological manipulation?” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, Y.-N. Chen, Eds. (Association for Computational Linguistics, 2024), pp. 17140–17161.
  • 67.D. Gurgurov et al., “Multilingual political views of large language models: Identification and steering” in Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, K. Inui et al., Eds. (The Association for Computational Linguistics, 2025), pp. 279–298.
  • 68.D. Ganguli et al., The capacity for moral self-correction in large language models. arXiv [Preprint] (2023). http://arxiv.org/abs/2302.07459 (Accessed 20 November 2025).
  • 69.Abdurahman S., et al. , Perils and opportunities in using large language models in psychological research. PNAS Nexus 3, pgae245 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Atari M., et al. , Pathogens are linked to human moral systems across time and space. Curr. Res. Ecol. Soc. Psychol. 3, 100060 (2022). [Google Scholar]
  • 71.Rai T. S., Fiske A. P., Moral psychology is relationship regulation: Moral motives for unity, hierarchy, equality, and proportionality. Psychol. Rev. 118, 57 (2011). [DOI] [PubMed] [Google Scholar]
  • 72.S. H. Schwartz, “Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries” in Advances in Experimental Social Psychology, M. P. Zanna, Ed. (Academic Press, 1992), vol. 25, pp. 1–65.
  • 73.Janoff-Bulman R., Carnes N. C., Surveying the moral landscape: Moral motives and group-based moralities. Pers. Soc. Psychol. Rev. 17, 219–236 (2013). [DOI] [PubMed] [Google Scholar]
  • 74.Schein C., Gray K., The theory of dyadic morality: Reinventing moral judgment by redefining harm. Pers. Soc. Psychol. Rev. 22, 32–70 (2018). [DOI] [PubMed] [Google Scholar]
  • 75.Dogruyol B., et al. , Validation of the moral foundations questionnaire-2 in the Turkish context: Exploring its relationship with moral behavior. Curr. Psychol. 43, 24438–24452 (2024). [Google Scholar]
  • 76.Winkelkotte F., Fobi D., Möhring M., Wild S., Testing psychometrics of the Moral Foundations Questionnaire-2 (MFQ-2) among pre-service teachers in Ghana. Curr. Psychol. 44, 6746–6759 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Thalmayer A. G., Toscanelli C., Arnett J. J., The neglected 95% revisited: Is American psychology becoming less American? Am. Psychol. 76, 116 (2021). [DOI] [PubMed] [Google Scholar]
  • 78.C. W. Leach, R. Bilali, S. Pagliaro, “Groups and morality” in APA Handbook of Personality and Social Psychology, Vol. 2: Group Processes, M. Mikulincer, P. R. Shaver, J. F. Dovidio, J. A. Simpson, Eds. (American Psychological Association, 2015), vol. 2, pp. 123–149.
  • 79.Frank M. C., Goodman N. D., Cognitive modeling using artificial intelligence. Annu. Rev. Psychol. 77, 543–566 (2025). [DOI] [PubMed] [Google Scholar]
  • 80.Dillion D., Tandon N., Gu Y., Gray K., Can ai language models replace human participants? Trends Cogn. Sci. 27, 597–600 (2023). [DOI] [PubMed] [Google Scholar]
  • 81.S. Havaldar et al., “Multilingual language models are not multicultural: A case study in emotion” in Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment &, Social Media Analysis, J. Barnes, O. De Clercq, R. Klinger, Eds. (Association for Computational Linguistics, Toronto, Canada, 2023), pp. 202–214.
  • 82.M. Mazeika et al., A standardized evaluation framework for automated red teaming and robust refusal. arXiv [Preprint] (2024). 10.48550/arXiv.2402.04249 (Accessed 20 November 2025). [DOI]
  • 83.A. Birhane et al., “The values encoded in machine learning research” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Association for Computing Machinery, 2022), pp. 173–184. [DOI] [PMC free article] [PubMed]
  • 84.Z. Kenton et al., Alignment of language agents. arXiv [Preprint] (2021). http://arxiv.org/abs/2103.14659 (Accessed 20 November 2025).
  • 85.Buick A., Copyright and AI training data-transparency to the rescue? J. Intellect. Prop. Law Pract. 20, 182–192 (2025). [Google Scholar]
  • 86.Y. Jernite, “Training data transparency in AI: Tools, trends, and policy recommendations” Hugging Face Blog (2023). https://huggingface.co/blog/yjernite/data-transparency. Accessed 5 February 2026.
  • 87.A. Zewail et al. , Moral stereotyping in large language models. OSF. https://osf.io/afq67/overview?view_only=64f2f6b78b6e4f0abb6e441e567b9ebe. Deposited 14 November 2025. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Data Availability Statement

Code data have been deposited in OSF (https://osf.io/afq67/?view_only=64f2f6b78b6e4f0abb6e441e567b9ebe) (87). All other data are included in the manuscript and/or SI Appendix.


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES