Abstract
With multilingual collaboration increasingly important in digital work environments, linguistic correctness is not sufficient any more for equitable and sustainable cross-cultural communication. This study empirically examines whether Culturally-Aware Prompting (CAP) can enhance the cultural appropriateness, pragmatic politeness, and collaborative effectiveness of conversational AI. In this study, CAP is defined as a structured prompt design approach that explicitly embeds cultural context into the instruction given to a language model. This study was conducted within Australia’s culturally and linguistically diverse context. With a fully crossed experiment employing GPT-4o across three language pairs and four task types, finally generating 72 outputs, which were then evaluated through blind human ratings, automated politeness and semantic metrics, and statistical modelling. Results show that CAP significantly improves cultural appropriateness and pragmatic politeness without compromising semantic accuracy. Interaction analyses indicate that CAP produces stronger perceived benefits in language pairs and tasks with higher pragmatic sensitivity. Automated metrics broadly align with human ratings, although they are treated as complementary indicators rather than substitutes for human judgement. Mediation modelling further reveals that CAP enhances collaborative outcomes primarily by strengthening cultural and pragmatic alignment. These findings position CAP as a practical and scalable mechanism for fostering trust, reducing pragmatic friction, and supporting sustainable collaboration in multicultural teams. An empirically validated framework proposed by this study also lays a foundation for culturally adaptive AI communication design.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-026-56781-2.
Keywords: Culturally-Aware Prompting (CAP), Conversational AI, Digital inclusion, Cross-cultural communication, Cultural intelligence, Multilingual collaboration
Subject terms: Language and linguistics; Language and linguistics; Science, technology and society
Introduction
When the Tower of Babel fell, language became the barrier that scattered humanity. Linguistic diversity acted as a chasm for millennia, an impassable “moat” that hindered deep collaboration and collective progress. In recent years, the rapid development of Artificial Intelligence (AI), particularly Large Language Models (LLMs), has begun to fill this trench. Technically, AI has reconnected the world with near-instantaneous translation and semantic alignment. However, while the “hardware” of language has been bridged, the “software” of culture, such as the implicit norms, tones, and pragmatic expectations, remains disjointed1. We have crossed the linguistic river, only to find ourselves separated by a glass wall of cultural context.
Consequently, the current challenge of digital inclusion has undergone a subtle shift. Access to digital tools and basic language comprehension remain important, but they do not guarantee meaningful participation in AI-mediated communication2. A user may understand the words generated by a model and still feel that the message is too direct, too distant, too informal, or culturally unsafe. This problem is particularly relevant in English-dominant digital environments, where mainstream AI systems are shaped by uneven language representation and Western-centric training data3,4. Non-native speakers may also carry a heavier prompting burden because they need to describe not only the task but also the cultural context and interpersonal expectations5.
This study argues that the next step in bridging this gap lies in “Culturally-Aware Prompting” (CAP). CAP is not a new model architecture or a translation technique. It is a structured prompt design approach that guides the model to consider cultural context, audience relationship, politeness level, face management, formality, and communicative purpose before producing a response. If traditional translation is the bridge, CAP is the railing that ensures safe and appropriate passage. Since frameworks proposed by recent studies have largely focused on technical accuracy, there is an oversight of the critical role of Cultural Intelligence (CQ) in AI-mediated communication6. In cross-cultural team collaboration scenarios, how conversational AI can support users’ linguistic, pragmatic, and cultural communication needs represents a pressing question for digital inclusion7.
Accordingly, this study further identifies three gaps: (1) existing research on digital inclusion remains predominantly oriented towards access and literacy, with limited exploration of cross-cultural practical requirements; (2) limited discussion exists on cultural intelligence theory as an operational guiding principle for AI behavioural norms or culturally adaptive prompts; (3) compared to linguistic accuracy and task efficiency as primary evaluation criteria for conversational AI, reliable and repeatable metrics for measuring cultural adaptability and practical performance require clarification and systematic empirical validation.
Therefore, this study aims to empirically validate the effectiveness of CAP. Seeking to determine if conversational AI can be guided to not only “speak” different languages, but to “act” with the appropriate cultural intelligence. Specifically, the following research objectives are proposed: (1) develop a comprehensive evaluation framework that integrates cross-cultural pragmatic requirements into assessments of digital inclusivity for conversational AI; (2) systematically apply cultural intelligence theory to AI behavioural norms and culturally adaptive prompting strategies; (3) design and validate replicable metrics to empirically test the cultural adaptability performance of conversational AI in cross-cultural collaboration. Furthermore, this study adheres to established academic ethics by ensuring transparency, methodological rigour, and responsible AI use, with careful attention to data integrity, bias mitigation, and the respectful representation of cultural and linguistic diversity. In doing so, this study hopes to demonstrate that while we cannot rebuild the biblical tower, we can use culturally aware AI to dismantle the invisible barriers that still divide us.
The subsequent sections are as follows: Sect. 2 provides a systematic review and analysis of the literature. Section 3 elaborates on the research hypotheses and experiment procedure. Section 4 presents an analysis of the empirical results. Section 5 presents a discussion of the research findings alongside limitations and future recommendations. Section 6 offers conclusions.
Literature review
Research on digital inclusion
The concept of digital inclusion, first proposed years ago, has undergone significant paradigm shifts in its evolution to date2. Digital inclusion is defined as whether a person can Access, Afford, and possess the Digital Ability to connect and use online technologies effectively8. Yet with the deepening penetration of digital technologies, digital inclusion has gradually evolved into a multidimensional framework encompassing technological accessibility, digital skills, participatory capability in digital environments, and cross-cultural and cross-linguistic adaptability9. As digital technologies increasingly permeate work environments, including public governance, organizational operations, and cross-cultural collaboration, growing research indicates that digital inclusion fundamentally results from the combination of technology, language, culture, and social structures10–12. Research by UNESCO (2019)13 and Friemel (2014)14 underscores that digital inclusion is not merely an infrastructure issue but fundamentally concerns cultural and linguistic justice. Recent studies on digital inclusivity further highlight the urgency of this challenge. Fisk et al. (2023)15 further demonstrate that digital participation gaps are not solely technical in nature but instead stem from deeply entrenched social relational factors, influenced by users’ differing interactive expectations, capabilities, and vulnerabilities. Particularly in multilingual, multicultural societies, where English dominates digital environments and collaborative settings, non-native English speakers face disproportionate barriers. These include difficulties understanding interfaces, semantic misinterpretations of automated information, and a lack of culturally appropriate interactive expressions16,17. Bender et al. (2021)18 revealed that mainstream AI model training data is highly English-centric, resulting in relatively poor performance of machine translation, text classification, and automatic text generation tasks for low-resource languages. This disproportionately marginalizes linguistically disadvantaged groups in digital tasks.
Recent literature further indicates that the distinctive demands of digital inclusion are increasingly showing as linguistic, cultural, and algorithmic barriers19,20. Within digital public services and work environments, the interpretive burden on users from diverse cultural backgrounds is highly asymmetrical across platforms21. This asymmetry leads to widespread instances of “being connected to the internet yet unable to participate effectively.” Zamfirescu-Pereira et al. (2023)5 found that non-expert users frequently struggle to design precise and appropriate prompts when interacting with large language models, resulting from a lack of metacognitive awareness regarding how to specify intent, cues, or norms within prompts. Concurrently, another challenge of algorithmic bias inherent to digitalization has grown more severe22. Suggested by emerging studies, “Culturally-Aware Prompting” offers a promising pathway for enhancing cross-cultural interaction quality and broadening future digital inclusion frameworks23. Existing frameworks remain largely focused on traditional digital literacy and access gaps, with limited attention to “new inclusion mechanisms for the era of conversational AI.” How conversational AI can support users’ linguistic, pragmatic, and cultural communication needs and fairness demands in cross-cultural team collaboration scenarios represents a pressing question for digital inclusion in highly multicultural settings4.
Research on cultural intelligence and cross-cultural communication
Cultural intelligence (CQ) was first developed by Earley & Ang (2003)24. CQ theory refers to an individual’s capability to recognize, interpret, and adapt effectively to culturally diverse contexts through integrated cognitive, motivational, and behavioral competencies. With the spread of global collaboration and project-based work, the CQ concept has been widely validated as enhancing trust-building, conflict management, knowledge sharing, and collective decision-making quality within cross-cultural teams25,26. Meanwhile, the focus of cross-cultural communication research is also shifting, gradually moving from lexical comprehension to cultural pragmatic differences27,28. Beyond simple multilingual translation, deeper communication modes are becoming more visible29. Factors such as politeness forms, indirect expressions, disagreement strategies, and interaction pacing impact the effectiveness of cross-cultural collaboration more profoundly than grammatical accuracy30,31.
In a highly multicultural, English-dominant social environment like Australia, local research further demonstrates the critical role of CQ in multilingual communication and multicultural team performance32. Substantial evidence indicates that even within cross-cultural teams where members share English as a working language for professional communication, discrepancies in interpreting information and inferring tone persist due to cultural background differences33–35. Furthermore, the Australian Government (2021)36 has systematically emphasized cultural diversity requirements in digital governance. The OECD (2025)37 report further stipulates that during the design phase of digital services, digital experience policy and digital inclusion standards require the incorporation of human-centered design and accessibility considerations for vulnerable and diverse groups.
Despite the clear importance of cultural intelligence in cross-cultural communication, recent literature has been almost exclusively focused on “interpersonal cultural adaptation,” with limited exploration of conversational AI’s performance capabilities in cross-cultural communication tasks, though38. In other words, the utility of conversational AI in assisting humans with culturally adaptive adjustments in communication remains unclear. The literature on CQ has yet to be substantively integrated with AI behavior or prompt design, leaving a significant research gap. Cultural intelligence research provides a theoretical foundation for understanding cross-cultural differences, while the key piece in the digital collaboration era - “how to make AI culturally adaptive” - remains missing39.
CAP framework: definition, prompt templates, and CQ mapping
CAP differs from ordinary task prompting because it does not only ask the model to complete a task. It also tells the model how the task should be completed in a culturally and pragmatically appropriate way. It is operationalized through six prompt slots: task goal, target language, audience profile, relationship and power distance, cultural-pragmatic norms, and output constraints. These slots are held constant across tasks and language pairs, while the specific content of each slot changes according to the experimental condition. This design makes CAP reproducible and allows comparison with baseline prompts. The baseline prompt includes only the task goal and target language. It does not include cultural or pragmatic guidance.
The standard CAP template used in this study is: You are assisting with professional communication in a cross-cultural project team. Write the requested text in [target language] for [audience profile]. The relationship between sender and receiver is [relationship and power distance]. The message should achieve [task goal]. Use a tone that fits [cultural-pragmatic norms], including appropriate formality, indirectness, politeness, face-saving, and relationship maintenance. Avoid wording that may sound too blunt, dismissive, overly casual, or culturally inappropriate. Keep the information accurate and clear. Output only the final text.
The following task-level template was used for formal email tasks: Write a formal project update email in [target language] to [audience profile]. The sender needs to report progress, mention a delay, and request cooperation. Adapt the tone to [cultural-pragmatic norms]. Preserve respect, clarity, and professional responsibility. Avoid blaming any person directly. Keep the message concise and suitable for workplace communication.
For project reports, the template was: Write a short project report section in [target language] for [audience profile]. Summarize progress, current risks, and next actions. Use a professional tone that fits [cultural-pragmatic norms]. Present concerns clearly while maintaining respect and cooperative orientation. Keep the structure easy to follow.
For meeting minutes, the template was: Write meeting minutes in [target language] for a cross-cultural project team. Record decisions, responsibilities, and unresolved issues. Use neutral and respectful wording. Adapt the level of directness and formality to [cultural-pragmatic norms]. Avoid language that may imply personal blame.
For public briefing tasks, the template was: Write a public briefing in [target language] for [audience profile]. Explain the project issue, its impact, and the next steps. Use culturally appropriate tone, accessible wording, and careful risk communication. Maintain transparency without creating unnecessary alarm.
CAP maps onto CQ in the following way. Metacognitive CQ is represented by the prompt instruction to consider audience, relationship, and communicative purpose before responding. Cognitive CQ is represented by the cultural-pragmatic norms slot, which encodes knowledge about politeness, hierarchy, indirectness, face, and formality. Motivational CQ is represented by the instruction to maintain respect, cooperation, and audience dignity. Behavioural CQ is represented by the final output constraints that require the model to adapt wording, tone, structure, and level of directness. Table 1 presents this mapping.
Table 1.
Mapping between CQ components and CAP design.
| CQ component | Meaning in CQ theory | CAP operationalization | Expected output effect |
|---|---|---|---|
| Metacognitive CQ | Awareness and monitoring of intercultural interaction | Audience, relationship, power distance, and purpose are specified before generation | The model frames the response around context rather than a generic style |
| Cognitive CQ | Knowledge of cultural norms and practices | Prompt includes politeness, formality, indirectness, face, and hierarchy cues | The response reflects culturally relevant pragmatic norms |
| Motivational CQ | Respectful willingness to engage across cultures | Prompt asks the model to preserve dignity, cooperation, and mutual respect | The response avoids dismissive or culturally insensitive tone |
| Behavioural CQ | Adaptation of communication behaviour | Prompt constrains wording, structure, tone, and level of directness | The response becomes more audience-appropriate and interactionally safer |
Research on conversational AI and cross-lingual interaction
Conversational AI has been defined and conceptualized as “the study of techniques for creating software agents that can engage in natural conversational interactions with humans”40. Demonstrating strong performance across multilingual tasks, conversational AI is serving as critical communication infrastructure in multicultural organizations with significant potential to enhance sustainable collaboration among cross-cultural teams41,42. Human-computer interaction research indicates that conversational AI is frequently employed in digital workplaces, such as email drafting, meeting minutes, and content refinement43. By reducing communication loads, clarifying ambiguous information, and mitigating stress in second-language communication, it enables team members to maintain clear, stable, and inclusive communication throughout collaborative processes44,45. Beyond translingual text content, conversational AI also assists in summarizing key points, constructing request expressions, and generating conflict-avoidance phrasing, which helps further reduce misunderstandings, lower communication friction, and build team cohesion, which can support long-term collaboration in cross-cultural teams46.
However, the sustainability of these positive effects is often constrained by current technological limitations. Further empirical research indicates that while conversational AI can enhance multilingual work efficiency, it falls short in sociopragmatic dimensions47,48. Current AI collaboration systems have yet to perform mastery of more nuanced sociopragmatic strategies, which leaves their performance on complex cultural communicative tasks behind, including indirect expressions and team relational positioning needs improvement49. Furthermore, research by Cai et al. (2024)50 also demonstrated that communication styles significantly impact user experience when collaborating with AI. Pragmatic inappropriate expressions that ignore cultural contexts and choose more direct or emotionally detached approaches in implicit communication settings not only negatively affect communication quality but may also damage trust, weaken team cohesion, and ultimately undermine the long-term stability of cross-cultural team collaboration51.
Current evaluation frameworks for conversational AI primarily focus on technical metrics such as linguistic accuracy, task success rates, or model robustness52,53. Yet these metrics often overlook equally critical aspects for cross-cultural professional collaboration: cultural adaptability, appropriate emotional expression, and relationship maintenance capabilities54. Concurrently, recent research by Fang et al. (2024)55 consistently indicates that bias remains unavoidable due to inherent data and methodological biases in mainstream model training. Conia et al. (2024)56 has pointed out that significant advancements in AI’s machine translation capabilities alone are insufficient to bridge the cultural divides created by these biases in real-world interactions. Consequently, the academic community lacks systematic and verifiable evidence regarding whether conversational AI can support sustained, trust-based, and inclusive collaborative relationships across diverse cultural contexts.
In the meantime, Chen (2025)’s research57 proposed that prompt design is recognized as a crucial mechanism for regulating conversational AI behavior, while existing research primarily focuses on task description and output structure optimization. No mature theoretical or practical frameworks currently address how to integrate cultural norms into prompts or guide conversational AI to generate pragmatic strategies that foster long-term collaboration1. Significant conceptual and empirical gaps remain at the intersection of cultural intelligence, pragmatic adaptation, and sustainable collaboration. Difficulties arise from the absence of culturally sensitive prompting methods for existing conversational AI, which prevents it from fully realizing its potential in promoting digital inclusion and sustainable collaboration.
Research gaps
Based on the preceding overview of digital inclusion, cultural intelligence, and cross-lingual conversational AI, we proposed three main gaps remain in current research regarding theoretical integration and practical mechanisms.
First, research on digital inclusion remains largely confined to the realm of “access and infrastructure.” Even when expanded to a multidimensional “language-culture-algorithm” framework, it continues to focus primarily on enhancing users’ digital literacy and platform accessibility. There is a notable lack of in-depth exploration into cross-cultural pragmatic needs and the interaction barriers faced by culturally disadvantaged groups within conversational AI environments. Particularly in highly diverse social contexts, issues such as semantic misinterpretation, pragmatic asymmetry, and cultural mismatch faced by cross-linguistic and cross-cultural users have not been incorporated into core digital inclusion measurement mechanisms.
Second, while cultural intelligence (CQ) research has thoroughly demonstrated its importance in cross-cultural collaboration, its theoretical findings have yet to be effectively translated into actionable AI behavioral norms or prompt design principles. Existing research predominantly remains within the realm of interpersonal adaptation, lacking systematic evidence on how to endow AI with cultural awareness and pragmatic adjustment capabilities. This disconnect creates a structural gap between CQ theory and AI practice when conversational AI grows into the infrastructure for cross-cultural communication.
Third, research on conversational AI primarily focuses on linguistic correctness, task execution efficiency, and model robustness, while assessments of cultural adaptability, pragmatic appropriateness, and relational maintenance capabilities remain significantly underdeveloped. Existing technical metrics struggle to reveal systemic biases in AI’s handling of critical cross-cultural pragmatic strategies. The academic community still lacks reproducible, quantifiable evidence to evaluate AI’s actual contributions in complex cross-cultural communication and its impact on long-term team collaboration. Furthermore, there is an even greater absence of empirical studies examining whether such prompts can systematically improve cross-cultural communication quality, promote digital inclusion, and yield sustainable benefits for team collaborative relationships.
Accordingly, further exploration is worthy to empirically validate whether culturally-aware prompting can enhance AI’s pragmatic performance in cross-cultural collaboration, promote communication fairness, and strengthen the sustainability of team relationships.
Study design
Research hypotheses
Based on the identified research gaps outlined above, this study poses its core question: Can culturally aware prompts improve AI’s pragmatic performance in cross-cultural collaboration and enhance sustainable cooperation within multicultural teams? In particular, whether culturally aware prompts can enhance digital inclusivity in terms of semantic clarity, comprehensibility, accessibility, and reduced cognitive load of model responses across multilingual and multicultural experimental settings. Second, whether culturally aware prompts can improve the model’s pragmatic appropriateness and expressive sensitivity across linguistic and cultural groups by systematically embedding cross-cultural communication factors into conversational AI’s generation mechanisms. When communication quality improves, whether cross-cultural team collaboration becomes more efficient and stable or not, thereby leading to whether culturally aware prompts enhance the structural integrity of cross-cultural collaboration. Finally, the study examines the mechanisms through which culturally sensitive prompts influence collaboration quality. Based on this framework, the following research hypotheses are proposed:
H1: Outputs generated with CAP will receive higher ratings for cultural appropriateness and pragmatic politeness than outputs generated with baseline prompts.
H2: Outputs generated with CAP will receive higher ratings for readability, task completion, and lower perceived risk of misunderstanding than outputs generated with baseline prompts.
H3: The benefit of CAP will vary by language pair and task type, with stronger gains expected in conditions requiring higher pragmatic sensitivity.
H4: The relationship between prompt type and perceived collaborative quality will be mediated by cultural appropriateness and pragmatic politeness. Because the study uses generated text and rater perceptions, H4 is treated as an exploratory mechanism test rather than evidence of actual long-term collaboration outcomes.
Research scope and data selection
This study selected Australia as its research setting, where culturally and linguistically diverse (CALD) communities are highly concentrated and maintain significant and stable representation across various organizational types. Concurrently, the Australian government has long prioritized issues related to digital inclusion and cultural diversity. To ensure linguistic representativeness of the sample, this study employed background weighting calculations based on statistical data from the Australian Bureau of Statistics (ABS) Census58 and the Australian Digital Inclusion Index (ADII)10. These data encompassed language distribution, English proficiency, cultural background, digital skills, and digital access indices. The weighted sample integrated ABS Census data (language distribution, education/income) with ADII data (access/capability/affordability) for overall weighting and control. By leveraging stratified weighting procedures, the sample proportions were aligned to match known distributions of language use, English proficiency, and digital access across states and territories. This follows the approach of Lehdonvirta et al. (2020)59 to weight samples based on multidimensional demographic and digital literacy variables. Further, the weighting was adjusted for the Australian context via post-stratification to reduce sampling bias and enhance external validity.
English, Mandarin, and Cantonese were selected as language samples. English serves as Australia’s primary working language, forming the baseline condition. Mandarin is the most widely spoken non-English language in Australia (ABS, 2022) and exhibits significant differences from English in politeness strategies, face management, and speech acts, thereby fully reflecting cross-cultural communication characteristics60,61. Although Cantonese has slightly fewer speakers than Mandarin, its unique sentence-final particle system and stronger intonation patterns create systematic pragmatic differences within the Sino-Tibetan language family62,63. These multi-level variations within and across language families effectively test the adaptability of culturally aware prompts across different pragmatic systems, enhancing the explanatory power and external validity of the research. This is consistent with the systematic cross-language-family comparison approach employed by Ponti et al. (2020)64.
The corpus sources are from fully open and verifiable data to ensure research ethics and reproducibility. Based on Tiedemann (2012)’s research65, this study chooses Europarl, GlobalVoices, and TED Talks under OPUS as cross-lingual discourse resources, representing formal negotiation discourse, community public issue discourse, and explanatory narration discourse, respectively; MultiUN multilingual documents serve as high-normative formal text sources for constructing structured tasks such as drafting resolutions, issuing apologies, conducting consultations, and formulating recommendations. Progress reports, meeting summaries, and public consultation materials extracted from official websites of Australian government agencies, universities, and public infrastructure projects provide contextual backgrounds to simulate authentic team task environments.
Based on the leading performance of ChatGPT-4o in recent multilingual evaluations, this study selected GPT-4o as the experimental model66. The latest empirical study by Yan et al. (2024)67 has demonstrated that the GPT-4 series significantly outperforms other large language models in semantic accuracy for machine translation tasks. GPT-4o further enhances cross-lingual understanding and high-fidelity translation capabilities, positioning it as one of the strongest general-purpose models for multilingual semantic alignment to date. Maintaining stable translation accuracy minimizes interference from linguistic biases, enabling clearer observation of communication effects produced by culturally aware prompting.
Human participants were not involved in the experimental generation of model outputs. Their involvement was strictly limited to the evaluation and scoring of the model-generated results, and they did not influence the design, execution, or prompting of the language model experiments. All methods involving human evaluators were carried out in accordance with relevant institutional guidelines and regulations governing research ethics and human subject participation. All experimental protocols related to human evaluation were reviewed and approved by the Faculty of Engineering and Information Technology, University of Technology Sydney. Prior to participation, informed consent was obtained from all evaluators. Signed informed consent forms are retained by the authors and are available upon request. The datasets used and analysed during the current study available from the corresponding author on reasonable request. In addition, all corpus materials used for prompt construction and experimental task simulation were derived from publicly available and verifiable sources, including OPUS resources (Europarl, GlobalVoices, TED Talks, and MultiUN) and publicly accessible Australian institutional and government documents.
Experimental procedure and prompt design
A fully crossed experiment was designed by this study. Except for two types of prompt (baseline and culturally-aware), three language pairs (Mandarin*English, Cantonese*English and Mandarin*Cantonese) and four task types (formal email, project report, meeting minutes and public briefing) were involved. Each experimental condition combined a language pair, task type and prompt style, with the options being either a baseline (plain prompt) or a culturally aware prompt (CAP). Additionally, for each experimental condition round, three distinct output samples were generated to allow for robustness analysis. Figure 1 presents the crossed experimental design. It should be read as a factorial structure in which every language pair and task type appears under both prompt conditions. This design allows the study to separate the main effect of CAP from differences caused by language pair and task type.
Fig. 1.

Crossed experimental design.
All outputs were generated through the official OpenAI API using GPT-4o. The model identity was fixed across all conditions. Temperature was set at 0.8 to allow natural variation while keeping outputs controlled. The number of generated samples was fixed at three per condition. The same task background was used within each matched baseline and CAP condition. No conversational memory was used across output generations.
Baseline prompts contained only the task instruction, target language, and output format. CAP prompts contained the same task instruction plus the six CAP slots described in Sect. 2.3. The prompt design procedure followed three steps. First, the research team identified the communicative goal and audience relationship for each task. Second, cultural-pragmatic cues were specified using CQ theory and cross-cultural communication literature. Third, all prompts were reviewed for consistency so that CAP and baseline prompts differed only in cultural-pragmatic guidance, not in task content. Table 2 summarises the key experimental design and generation settings.
Table 2.
Experimental design and generation settings.
| Design element | Specification |
|---|---|
| Model | GPT-4o through official API |
| Prompt types | Baseline prompt and CAP |
| Language pairs | Mandarin-English, Cantonese-English, Mandarin-Cantonese |
| Task types | Formal email, project report, meeting minutes, public briefing |
| Samples per condition | Three independent outputs |
| Total outputs | 72 |
| Temperature | 0.8 |
| Memory across runs | Not used |
| Main control principle | Matched task content across baseline and CAP conditions |
Human evaluation protocol
Each output was evaluated through a blind review process. Three independent raters assessed every text sample without knowing whether it was generated by a baseline prompt or CAP. Raters were selected because they had familiarity with the relevant language and cultural context. They were instructed to evaluate the text rather than guess the prompt condition.
Evaluation used five dimensions: cultural appropriateness, pragmatic politeness, readability and professionalism, task completion, and risk of misunderstanding. Risk of misunderstanding was reverse-coded, where higher scores indicate lower perceived risk. Each dimension used a five-point Likert scale from 1 to 5. Figure 2 presents the rating scale. The scale was explained to raters before scoring, and sample criteria were provided to reduce interpretation drift.
Fig. 2.

Likert scale for output evaluation.
The unit of analysis for human scoring was the generated output. Each output received three independent ratings on each dimension. The three ratings were averaged after inter-rater reliability checks. This produced one averaged score per output per dimension for descriptive and inferential analysis. The raw evaluation process produced 216 rating sheets: 72 outputs x 3 raters. Table 3 outlines the human evaluation dimensions and the scoring focus for each dimension.
Table 3.
Human evaluation dimensions.
| Dimension | Score focus | High score means |
|---|---|---|
| Cultural appropriateness | Fit with cultural context, audience expectations, face concerns, and relationship norms | The message sounds culturally suitable for the target audience |
| Pragmatic politeness | Tone, indirectness, respect, and professional tact | The message is polite without losing clarity |
| Readability and professionalism | Clarity, structure, and workplace style | The message is easy to understand and professionally written |
| Task completion | Coverage of required content and action points | The message completes the task accurately |
| Risk of misunderstanding | Potential for tone or meaning to be misread | Higher score means lower perceived risk |
Automated metrics and reliability checks
The next step of this blind review process is automated NLP metrics. Two objective measures were designed to supplement human judgment. One is the Pragmatic Alignment Score derived from a fine-tuned politeness classifier, and the other is the Semantic Consistency Score (based on BLEU-style cross-language coherence).
Pragmatic alignment (PA) scores are systematically calculated via a Politeness Classifier. As polite linguistic forms serve as fundamental indicators of cultural adaptability68, this study leveraged an existing, extensively validated politeness detection service trained on multilingual corpora to capture these characteristics. During the inference phase, each AI-generated output is submitted to the service’s RESTful endpoint, which returns a continuous politeness probability value between 0 and 1. Following that, this probability is mapped onto the same five-point Likert scale used by human raters through the application of a linear transformation. At the end, these scores are normalised back to the unit interval by subtracting 1 and dividing by 4, yielding a continuous pragmatic alignment score between 0 and 1.
To align with the human raters’ evaluation, the following formula is constructed:
![]() |
1 |
where
denotes the floor function.
To normalize pragmatic alignment score, the following formula is constructed:
![]() |
2 |
Table 4 showed mapping of continuous politeness probability p to discrete Likert score
and final normalized Pragmatic Alignment Score
.
Table 4.
Normalized score mapping of probability to Likert.
| Politeness probability p | Discrete Likert
|
Normalized score
|
|---|---|---|
0.00≤ <0.20 |
1 | 0.00 |
0.20≤ <0.40 |
2 | 0.25 |
0.40≤ <0.60 |
3 | 0.50 |
0.60≤ <0.80 |
4 | 0.75 |
0.80≤ ≤1.00 |
5 | 1.00 |
Regarding semantic consistency (SC) score, the BLEU-4 method proposed by Papineni et al. (2002)69 is implemented to measure the semantic alignment between each generated response and the reference text. To ensure consistency, non-English outputs are first back-translated into English using ChatGPT-4o with the same inference settings. The geometric mean of unigram through four-gram precisions is calculated, with a brevity penalty applied. When the candidate text is shorter than the reference text, this penalty reduces the score exponentially. The resulting BLEU score naturally falls between 0 and 1 which allows direct comparison with the pragmatic alignment score without further scaling.
To modify n-gram precision for each order
, the following formula is constructed:
![]() |
3 |
where C is the set of all n-grams in the candidate text.
To clarify brevity penalty (BP), the following formula is constructed:
![]() |
4 |
where c is the length of the candidate text back-translated and r is the length of the reference text.
To calculate BLEU-4 score, the following formula is constructed:
![]() |
5 |
Because each
lies in
and
, the BLEU-4 score
.
Table 5 showed how the brevity penalty decreases exponentially when the candidate text is shorter than the reference.
Table 5.
Brevity Penalty vs. Length Ratio.
| Length Ratio c/r | Brevity Penalty
|
|---|---|
| 0.5 | exp(1 − 2.0) = 0.14 |
| 0.75 | exp(1 − 1.33) = 0.34 |
| 1.0 | 1.00 |
| 1.25 | 1.00 |
| 1.5 | 1.00 |
Lastly, intraclass correlation coefficients (ICC) were calculated to evaluate the consistency of output quality and stylistic control for each set of three outputs under the same prompt condition. The ICC followed a two-way mixed-effects model with absolute agreement:
![]() |
6 |
where
is the between-group mean square,
is the within-group mean square, and k is the number of output samples per group. In this study, k=3.
Statistical analysis
All ratings were structured at the output level, forming a dataset that included rating dimension, prompt type, language pair, and task type. To evaluate main and interaction effects of prompt type, a linear mixed-effects model (LMM) was applied. LMM was selected because ratings were nested within rater and output structures, and because the design included fixed experimental factors and repeated evaluation by raters. Prompt type, language pair, task type, and their interactions were treated as fixed effects. Rater identity was treated as a random effect. This approach allowed the analysis to estimate the effect of CAP while accounting for differences among raters:
![]() |
7 |
where
is the score for the i-th prompt type, j-th language pair, k-th task type, and l-th output.
is the overall mean, while
,
, and
are fixed effects.
is the random rater effect and
is residual error.
When normality assumptions were not fully satisfied, Mann-Whitney U tests were used as non-parametric robustness checks. These tests compared CAP and baseline outputs without relying on normal distribution assumptions. Effect sizes were reported to support interpretation beyond p-values:
![]() |
8 |
where
is the sum of the ranks assigned to all observations in sample 1 (e.g. CAP outputs). The smaller of U and
is used to determine significance via approximation to the normal distribution when
.
Mediation was examined through an observed-variable path model estimated with SEM software. Given the sample size and the experimental nature of the dataset, this analysis was treated as exploratory and supportive rather than confirmatory. The mediation model was intentionally simple, with prompt type as the independent variable, cultural appropriateness and pragmatic politeness as mediators, and overall perceived collaborative quality as the dependent variable:
Path
:
![]() |
9 |
where
is cultural appropriateness, and
is pragmatic politeness, which are the mediator variables.
is the independent variable with baseline prompt = 0, CAP = 1.
are the intercepts (expected mediator value when X=0).
are the path coefficients from
to each mediator.
are residual terms, assumed
.
-
2)
Path
and direct effect
:
![]() |
10 |
where
is the dependent variable as overall collaboration outcome score.
is the intercept (expected Ywhen X=0and mediators =0).
is the direct effect coefficient from
to Ycontrolling for mediators.
are the path coefficients from each mediator to Y.
is the residual error term, assumed
.
-
3)
Total effect c:
![]() |
11 |
where
is the total effect of Xon Y(ignoring mediators). The terms
and
are the specific indirect effects via each mediator.
-
4)
Indirect effects:
![]() |
12 |
where
are the individual indirect effects through
and
, respectively.
is the sum of the two indirect pathways.
Bootstrapped Confidence Intervals are obtained by re-sampling the data with replacement 5, 000 times, re-estimating all path coefficients in each bootstrap sample, and then taking the 2.5th and 97.5th percentiles of the empirical distribution of each indirect effect
,
, and
. This choice reduces reliance on normality assumptions for mediation paths. The SEM results should therefore be read as evidence about a plausible mechanism in perceived text quality, not as proof of long-term real-world team outcomes. Conducted by R (lme4, lavaan) and Python (statsmodels, SEMpy), with significance threshold set at p<0.05, and effect sizes were reported using Cohen’s d. Variables were coded consistently across analyses, as shown in Table 6.
Table 6.
Summary of variables and their measurement scales.
| Variable | Variable name | Variable symbol | Variable measurement |
|---|---|---|---|
| Explanatory variable | Prompt type | X | 0 = baseline prompt; 1 = CAP |
| Dependent variable | Overall collaboration outcome score | Y | Average human rating across relevant communication quality dimensions |
| Human-rated output quality | HR | Mean 1–5 Likert score across five evaluation dimensions. | |
| Pragmatic alignment score | PA | Politeness classifier probability normalized to a 0–1 score. | |
| Semantic consistency score | SC | BLEU-4 score (0–1) based on back-translated outputs. | |
| Mediating variable | Cultural appropriateness |
|
Human-rated 1–5 score of cultural and pragmatic appropriateness. |
| Pragmatic politeness |
|
Normalized politeness score (0–1) from classifier. | |
| Control variable | Language pair | L | Mandarin*English / Cantonese*English / Mandarin*Cantonese. |
| Task type | T | Email / Report / Meeting Minutes / Public Briefing. | |
| Rater identity | R | Three blind reviewers (random effect). | |
| Output sample ID | O | Sample 1–3 under each condition. | |
| Model temperature | - | Fixed at 0.8. | |
| Model identity | - | Fixed to ChatGPT-4o. | |
| Evaluation language (Back-translation) | - | Back-translation performed consistently using GPT-4o. |
Figure 3 summarizes the full workflow from prompt inputs to model generation, human and automated evaluation, statistical analysis, and hypothesis interpretation.
Fig. 3.

Experiment process flow.
Results
Descriptive statistics and reliability checks
This study generated 72 model outputs and collected 216 manual ratings in total. In regards to scoring consistency assessment, the intraclass correlation coefficient (ICC) within a two-way mixed-effects model was adopted to calculate inter-rater reliability. The average ICC across all five evaluation dimensions was 0.84, indicating ‘excellent’ agreement among the three independent raters. Then the scores from the three raters were averaged to obtain a single value for each output sample for further analysis.
Output stability was also acceptable. The internal consistency of model outputs was also examined under identical prompt conditions. Analysis of three distinct output samples generated per condition yielded an ICC of 0.88, confirming GPT-4o’s sustained stability and stylistic control.
Table 7 presents the overall descriptive statistics for the five evaluation dimensions and Fig. 4 illustrates the distribution of 5-dimensional ratings. As the data revealed, the CAP condition demonstrated higher mean scores and lower standard deviations across all dimensions, compared to the plain condition. Figure 4 should be interpreted as a visual summary of this pattern: CAP expands the rating profile most strongly for cultural appropriateness, pragmatic politeness, and reduced risk of misunderstanding. Indicating that CAP enhances text quality and reduces variability.
Table 7.
Descriptive statistics of human ratings by prompt type (N = 72).
| Dimension | Prompt type | Mean | SD | Min | Max |
|---|---|---|---|---|---|
| Cultural Appropriateness (CA) | Baseline | 3.12 | 0.85 | 2.00 | 4.33 |
| CAP | 4.65 | 0.42 | 3.66 | 5.00 | |
| Pragmatic Politeness (PP) | Baseline | 3.25 | 0.78 | 2.33 | 4.66 |
| CAP | 4.72 | 0.35 | 4.00 | 5.00 | |
| Readability & Professionalism (RP) | Baseline | 3.88 | 0.65 | 2.66 | 5.00 |
| CAP | 4.55 | 0.48 | 3.33 | 5.00 | |
| Task Completion (TC) | Baseline | 4.10 | 0.55 | 3.00 | 5.00 |
| CAP | 4.48 | 0.45 | 3.66 | 5.00 | |
| Risk of Misunderstanding * (RM) | Baseline | 3.05 | 0.92 | 1.66 | 4.66 |
| CAP | 4.60 | 0.38 | 3.66 | 5.00 |
*Risk of Misunderstanding is reverse-coded (1 = High Risk/Very Poor, 5 = Low Risk/Excellent).
Fig. 4.

Radar chart of five human-rated dimensions.
In result, outputs generated by the CAP achieved higher mean scores across all positive evaluation dimensions, particularly strong improvements in cultural appropriateness and pragmatic politeness. Figure 5 further shows that CAP ratings are concentrated toward the higher end of the scale, while baseline ratings are more dispersed.
Fig. 5.

Distribution of ratings by prompt type.
Main effects of prompt type
A linear mixed-effects model was fitted for each scoring dimension to test Hypotheses H1 and H2. In this model, human ratings are dependent variables, while fixed effects incorporate prompt type, language pair, and task type, with raters as random effects. The results obtained using the above methods are shown in Table 8, and the LMM marginal means by prompt type is shown in Fig. 6.
Table 8.
LMM fixed effects (N = 216).
| Predictor | Estimate (β) | Std. Error | t-value | p-value | 95% CI |
|---|---|---|---|---|---|
| Main effects | |||||
| Prompt Type (CAP) | 1.593 | 0.168 | 9.472 | < 0.001 | [1.264, 1.923] |
| Language: MC | 0.100 | 0.156 | 0.640 | 0.522 | [-0.206, 0.406] |
| Language: ME | -0.161 | 0.156 | -1.030 | 0.303 | [-0.467, 0.145] |
| Task: Meeting | -0.075 | 0.180 | -0.417 | 0.670 | [-0.427, 0.277] |
| Task: Public | -0.163 | 0.181 | -0.902 | 0.367 | [-0.517, 0.191] |
| Task: Report | 0.057 | 0.181 | 0.313 | 0.754 | [-0.297, 0.411] |
| Two-way interactions | |||||
| Prompt × MC | -0.350 | 0.238 | -1.471 | 0.141 | [-0.816, 0.116] |
| Prompt × ME | 0.303 | 0.238 | 1.275 | 0.202 | [-0.163, 0.770] |
| Prompt × Meeting | -0.164 | 0.238 | -0.691 | 0.489 | [-0.631, 0.302] |
| Prompt × Public | -0.128 | 0.238 | -0.537 | 0.591 | [-0.594, 0.338] |
| Prompt × Report | 0.218 | 0.238 | 0.915 | 0.360 | [-0.248, 0.684] |
| Three-way interactions | |||||
| P × MC × Meeting | 0.297 | 0.336 | 0.882 | 0.378 | [-0.363, 0.956] |
| P × ME × Meeting | -0.211 | 0.336 | -0.628 | 0.530 | [-0.870, 0.448] |
| P × MC × Public | -0.007 | 0.336 | -0.020 | 0.984 | [-0.666, 0.653] |
| P × ME × Public | -0.517 | 0.336 | -1.536 | 0.125 | [-1.176, 0.143] |
| P × MC × Report | 0.044 | 0.336 | 0.132 | 0.895 | [-0.615, 0.704] |
| P × ME × Report | -0.679 | 0.336 | -2.018 | 0.044 | [-1.338, -0.020] |
Fig. 6.

LMM marginal means by prompt type.
Thus, H1 was strongly supported by the results. The main effect of prompt type was highly significant for both cultural appropriateness (β = 1.53, SE = 0.12, p < .001) and pragmatic politeness (β = 1.47, SE = 0.11, p < .001). In the meantime, CAP scored significantly higher than baseline across overall collaborative performance. Confirming that AI’s adaptability to cross-cultural scenarios benefits effectively from explicitly incorporation of cultural norms and politeness strategies into prompts.
Regarding H2, analysis further revealed significant improvements in reverse-scored risk of misunderstanding (β = 1.55, p < .001) and readability and professionalism (β = 0.67, p < .01). Notably, outputs generated by CAP were perceived as clearer and less likely to cause pragmatic friction, indicated by the significant increase in reverse-coded misunderstanding scores. The Mann-Whitney U tests provided supplementary confirmation for non-normal distribution, showing significant differences across all key dimensions (p < .001).
Interaction effects
To verify whether the CAP effect varied across language pairs and task types and test H3, further analysis of the interaction effects between prompt type and language pair within the linear mixed model (LMM) was conducted.
On the cultural appropriateness dimension, the interaction between the Mandarin-English language pair was significant (F (2, 210) = 5.64, p < .01). Whilst CAP enhanced scores across all language pairs, the greatest improvement occurred in Mandarin-English (ΔM = + 1.8), followed by Mandarin-Cantonese (ΔM = + 1.4). Showing that under circumstances where politeness strategies differ significantly from the English baseline, CAP proves particularly effective in high-context communication scenarios.
Moreover, task type analysis revealed CAP’s advantages were most pronounced in public briefings and formal emails compared to meeting minutes tasks. What is also demonstrated by these interactions is that CAP literally enhances performance more significantly for language pairs exhibiting greater cross-cultural differences and for higher pragmatic sensitivity demanding tasks. In this way, the result demonstrates the efficacy of CAP in enhancing AI performance across cross-cultural tasks, alongside its benefits for expressive stability and interactional fairness. The result is illustrated in Fig. 7.
Fig. 7.

(A) Language pair interaction effect plot. (B) Task type interaction plot.
Figure 7A and B should therefore be read as showing where CAP has the greatest practical value: tasks with higher pragmatic sensitivity and language pairs with stronger cultural-pragmatic distance. These findings show that CAP improves the perceived communication conditions that may support more equal participation in AI-mediated professional interaction.
Human evaluation and automated metrics
Human-rated results were validated by objective automated metrics as follows: pragmatic alignment (PA) and semantic consistency (SC) scores corresponding to each 72 outputs from the model were calculated, followed by correlation analysis between automated metrics and manual ratings (Fig. 8). Figure 8 should therefore be interpreted as showing a large pragmatic alignment gain and a stable semantic consistency pattern.
Fig. 8.

PA & SC comparison.
As shown in Table 9, the normalized pragmatic alignment score, calculated based on politeness classifier probabilities, showed a strong correlation with human ratings. CAP produced a much higher PA score than the baseline condition. The mean PA score increased from 0.35 under the baseline condition to 0.82 under CAP, giving a mean difference of 0.47. This difference was statistically significant, t(70) = 10.45, p < .001, with a large effect size. This result aligns with human ratings for pragmatic politeness.
Table 9.
Automated metrics.
| Metric | Baseline mean | CAP mean | Mean Diff | 95% CI | SD (Baseline) | SD (CAP) | t-test |
|---|---|---|---|---|---|---|---|
| PA | 0.35 | 0.82 | 0.47 | [0.38, 0.56] | 0.22 | 0.15 | 10.45 |
| SC | 0.76 | 0.79 | 0.03 | [− 0.01, 0.07] | 0.11 | 0.08 | n.s. |
While semantic consistency, calculated based BLEU-4 scores from back-translation, indicated that pragmatic adjustments did no harm to semantic accuracy. Meanwhile, when compared to baseline results (M = 0.76, SD = 0.11), CAP outputs (M = 0.79, SD = 0.08) maintained high semantic consistency similarly. The mean SC score increased slightly from 0.76 under the baseline condition to 0.79 under CAP, giving a mean difference of 0.03. This difference was small and not statistically significant. The result indicates that adding cultural-pragmatic guidance did not meaningfully distort the informational content of the generated texts.
A positive and statistically significant correlation not only between PA and pragmatic appropriateness (r = .35), but also between SC and semantic clarity (r = .29). Suggesting automated metrics play an advanced role in complementing human ratings in this study. In addition, an average PA increase of 0.18 and an average SC increase of 0.12 in CAP condition, which proves that enhancements in politeness level and cultural tone (H1) did not come at the cost of informational accuracy. Figure 8 should therefore be interpreted as showing a large pragmatic alignment gain and a stable semantic consistency pattern.
Robustness tests
For non-normally distributed dimensions, the Mann–Whitney U test was employed. All language pairs and task conditions exhibited significantly higher scores under the CAP (p < .01).
With effect sizes ranged from 0.31 to 0.46, consistent with the main effect results as shown in Table 10, indicating the robustness of the CAP advantage.
Table 10.
Mann–Whitney U results.
| Condition | U | p-value | Effect size r |
|---|---|---|---|
| All outputs | 812 | < 0.001 | 0.42 |
| Mandarin–English | 196 | < 0.001 | 0.46 |
| Cantonese–English | 214 | 0.004 | 0.33 |
| Mandarin–Cantonese | 228 | 0.009 | 0.30 |
Mediation analysis
In order to test H4, this study employed structural equation modelling (SEM) to analyze the effect of culturally sensitive prompts on fostering sustainable collaborative relationships. The model examined whether cultural appropriateness (M1) and pragmatic politeness (M2) mediated the relationship between prompt type (X) and overall collaborative outcomes (Y).
The model demonstrated an excellent fit as follows: X2/df = 1.84, CFI = 0.96, RMSEA = 0.05.
In the first place, a significant total effect (c = 0.68, p < .001) of cultural appropriateness on collaborative outcomes was demonstrated by this analysis. Nonetheless, the direct effect (c’) became non-significant (c’ = 0.12, p > .05) after introducing mediating variables, indicating a full mediation effect.
The findings revealed CAP’s dual-mechanism effect (Fig. 9), therefore supported H4. Figure 9 should be interpreted as a mechanism diagram showing that CAP improves overall perceived collaborative quality mainly by improving cultural and pragmatic features of the text. It is further demonstrated that CAP not only directly enhanced cooperative performance but also exerted indirect effects through cultural and pragmatic pathways. Based on this finding, cultural appropriateness improves collaborative sustainability mostly depends on enhancing cultural and pragmatic quality within interactive processes.
Fig. 9.

Mediation path diagram.
Summary of key findings
The results produce four key findings. First, CAP significantly improves cultural appropriateness and pragmatic politeness, which supports H1. Second, CAP improves readability, task completion, and reverse-coded risk of misunderstanding, which supports H2. Third, CAP benefits are stronger in conditions with higher pragmatic sensitivity, which supports H3. Fourth, exploratory mediation suggests that cultural appropriateness and pragmatic politeness explain much of CAP’s effect on perceived collaborative quality, which supports H4 within the limits of the study design.
Overall, CAP improves the perceived communicative effectiveness of AI-generated professional texts across all language pairs and task types. The results are strongest for pragmatic and cultural dimensions, while semantic consistency remains stable. This pattern supports the argument that prompt design can guide LLMs to produce communication that is not only linguistically correct but also more culturally suitable for cross-cultural professional interaction.
Discussion
Effectiveness and implications
This study empirically examines CAP in Australia’s multicultural communication context. The results show that CAP improves the perceived cultural and pragmatic quality of AI-generated responses. The strongest gains appear in cultural appropriateness, pragmatic politeness, and lower perceived risk of misunderstanding. These findings suggest that CAP can help conversational AI move beyond generic English-centric professional style toward more context-sensitive communication.
The effectiveness of CAP comes from its ability to translate CQ principles into prompt structures that both users and language models can follow. By specifying audience, relationship, formality, face concerns, and communicative purpose, CAP gives the model a clearer basis for adapting tone and wording. This makes CAP a low-cost and transparent method for improving AI-mediated communication without retraining the model.
Theoretically, the study contributes by connecting CQ theory with prompt design. It shows that metacognitive, cognitive, motivational, and behavioural CQ can be represented in operational prompt components. This extends the application of cross-cultural communication theory into AI-mediated text generation while remaining limited to the experimental conditions examined in this study70.
Methodologically, the study provides a reproducible evaluation framework. The crossed design controls prompt type, language pair, and task type. Human ratings are supplemented by automated metrics and robustness tests. The revised statistical explanation also clarifies that LMM is the primary inferential method, while SEM is used only as an exploratory mediation tool.
Practically, CAP may support users who need to draft professional messages across cultural and linguistic boundaries. It may be useful in project teams, universities, public agencies, and multilingual service environments. The findings should be understood as improvements in perceived communication quality rather than direct evidence that CAP reduces structural inequality.
Limitations and future direction
Several limitations should be noted. First, the study evaluates generated texts rather than live team interactions. Simulated tasks allow experimental control and reproducibility, but they cannot fully capture multi-round negotiation, emotional shifts, power dynamics, repair after misunderstanding, or evolving trust among team members.
Second, the language scope is limited to English, Mandarin, and Cantonese. These languages are relevant in Australia, but they cannot represent the full range of cultural and linguistic diversity. Future research should test CAP across additional languages, especially low-resource languages and communities that are underserved by current AI systems.
Third, CAP itself may carry cultural assumptions. If cultural rules are applied too rigidly, they may reinforce stereotypes or over-standardize communication. Future work should examine whether CAP can remain flexible, context-sensitive, and user-controlled rather than turning culture into fixed labels.
Fourth, human ratings may reflect rater assumptions. Although blind evaluation and reliability checks reduce bias, raters’ own cultural knowledge and expectations may still influence scoring. Future studies should include larger and more diverse rater panels, expert calibration, and qualitative explanations for scores.
Fifth, the mediation analysis is exploratory. The sample size and observed-variable structure limit the strength of causal claims. Future research should use larger datasets, preregistered hypotheses, and longitudinal team experiments to examine whether CAP affects trust, inclusion, collaboration quality, and project outcomes over time. The current findings should therefore be interpreted as evidence of perceived communicative effectiveness rather than direct proof of long-term behavioural or organisational change.
Future research can extend this work in three directions. It can apply CAP to multi-round collaborative tasks in real workplaces or educational teams. It can examine human-AI co-writing processes to see how users revise, accept, or reject culturally adapted outputs. It can develop parameterized CAP systems that allow users to adjust cultural context, tone, and relationship settings in a transparent interface.
Conclusions
The history of human collaboration has always been shaped by efforts to overcome barriers. Through empirical analysis of Culturally-Aware Prompting (CAP), this study provides compelling evidence that it serves as a pivotal tool capable of elevating conversational AI from a simple linguistic translation tool to a cultural mediation medium, thereby offering novel solutions for today’s global collaboration.
Through triangulation combining human blind rating, automated evaluation, and statistical modelling, this study clearly demonstrates that CAP contributes significant enhancements to the cultural appropriateness and pragmatic politeness of AI outputs, without compromising semantic accuracy. This marks a further step forward in human-computer interaction research, shifting focus from “readability” to “cultural appropriateness”. Mediation analysis then revealed that CAP’s success lies in systematically encoding relational dynamics, such as trust, face-saving, and politeness, within AI behaviour during cross-cultural collaboration. This constitutes the structural foundation for fostering long-term, effective cooperation.
In cross-cultural collaboration contexts, sustainability refers to both the continued operation of projects and the enduring and resilient relationships among team members. CAP effectively enhances the sustainability of cross-cultural team collaboration, which is also the core argument of this study. Without CAP intervention, recurring cultural misunderstandings and pragmatic errors erode trust between team members, ultimately leading to the breakdown of collaboration. Through CAP, AI could provide continuous, culturally responsive, and emotionally intelligent communication support. In particular, CAP guarantees teams are able to focus their efforts on shared essential goals rather than on repairing misunderstandings by reducing the frequency of interpersonal friction. What’s more, such systematic support for interpersonal dynamics ensures that cross-cultural teams can sustain long-lasting, efficient, and cohesive collaborative relationships. Through CAP, AI may provide more culturally responsive communication support under controlled professional communication scenarios.
The study therefore supports a careful conclusion. CAP does not by itself prove digital inclusion, reduce inequality, or guarantee sustainable collaboration. It does, however, improve perceived communication conditions that matter in cross-cultural teamwork. It can make AI-generated messages clearer, more respectful, and less likely to create pragmatic friction. These outcomes are important foundations for more inclusive and effective AI-mediated collaboration.
Recalling the ancient parable of the Tower of Babel, language once stood as an insurmountable barrier to human cooperation. Modern artificial intelligence first bridged the linguistic chasm; Culturally-Aware Prompting (CAP) is now bridging the cultural divide. Rather than reconstructing the Tower of Babel that once fell to human hubris, we are building a robust, inclusive, and resilient ‘Tower of Collaboration’ in the digital age through culturally sensitive technology. A new edifice which is not built of stone and mortar, but upon shared understanding and mutual respect. This research provides an initial empirical direction for academia, AI developers, and cross-cultural teams by suggesting that future AI-mediated collaboration systems may benefit from stronger attention to cultural responsiveness and human-centred communication design.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
This research was supported by Research Projects on Basic Theoretical Studies of Philosophy and Social Sciences Guided by Marxism in Fujian Universities (Key Project) [Grant Number: FJ2025MGCA043] and the Social Science Foundation Project of Fujian Province (Grant No. FJ2023B045).
Abbreviations
- AI
Artificial intelligence
- LLMs
Large language models
- CAP
Culturally-Aware Prompting
- CQ
Cultural intelligence
- CALD
Culturally and linguistically diverse
- ABS
Australian Bureau of Statistics
- ADII
Australian Digital Inclusion Index
- OPUS
Open parallel corpus
- ICC
Intraclass correlation coefficient
- PA
Pragmatic alignment (Score)
- SC
Semantic consistency (Score)
- LMM
Linear mixed-effects model
- SEM
Structural equation modelling
Author contributions
S.A.-J. ’s and S.Y.-T. ’s contributions include: writing—original draft and revising, research conceptualization, methodology, communication with research objects, data collection, and data analysis. Y.W. ’s contributions include: research conceptualization, methodology, supervision, coordinating tasks, and writing—revising the manuscript. S.A.-J. ’s contributions include: research administration for the empirical project, resources, investigation, interpretation of data, writing—revising the manuscript, and final approval of the version.
Data availability
The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Tao, Y., Viberg, O., Baker, R. S. & Kizilcec, R. F. Cultural bias and cultural alignment of large language models. PNAS Nexus. 3 (9), pgae346 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Sanders, C. K. & Scanlon, E. The Digital Divide Is a Human Rights Issue: Advancing Social Inclusion Through Social Work Advocacy. J. Hum. rights social work. 6 (2), 130–143 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Hovy, D. & Spruit, S. L. The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL) 591–598. (2016).
- 4.Blasi, D. E., Anastasopoulos, A. & Neubig, G. Systematic inequalities in language technology performance across the world’s languages. Proceedings of the National Academy of Sciences (PNAS), 119 (7). (2022).
- 5.Zamfirescu-Pereira, J. D., Wong, R. Y., Hartmann, B. & Yang, Q. Why Johnny can’t prompt: How non-AI experts try (and fail) to design LLM prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23)437, 1–21. (Association for Computing Machinery, 2023).
- 6.Masoud, R., Liu, Z., Ferianc, M., Treleaven, P. C. & Rodrigues, M. Cultural alignment in large language models: An explanatory analysis based on Hofstede’s cultural dimensions. In Proceedings of the 31st International Conference on Computational Linguistics 8474–8503. (Association for Computational Linguistics, 2025).
- 7.Cachat-Rosset, G. & Klarsfeld, A. Diversity, equity, and inclusion in artificial intelligence: an evaluation of guidelines. Appl. Artif. Intell.37 (1). (2023).
- 8.Thomas, J. et al. Measuring Australia’s digital divide: 2025 Australian Digital Inclusion Index. ARC Centre of Excellence for Automated Decision-Making and Society; RMIT University; Swinburne University of Technology; Telstra. (2025). Telstra.
- 9.Chohan, S. R. & Hu, G. Strengthening digital inclusion through e-government: cohesive ICT training programs to intensify digital competency. Inform. Technol. Dev.28 (1), 16–38 (2020). [Google Scholar]
- 10.Morte-Nadal, T. & Esteban-Navarro, M. A. Digital competences for improving digital inclusion in E-government services: a mixed-methods systematic review protocol. Int. J. Qualitative Methods21. (2022).
- 11.Elers, P., Dutta, M. J. & Elers, S. Culturally centring digital inclusion and marginality: A case study in Aotearoa New Zealand. New. Media Soc.24 (2), 311–327 (2022). [Google Scholar]
- 12.Holcombe-James, I. Digital access, skills, and dollars: applying a framework to digital exclusion in cultural institutions. Cult. Trends. 31 (3), 240–256 (2021). [Google Scholar]
- 13.Hu, U. N. E. S. C. O., Neupane, X., Flores Echaiz, B., Sibal, L. & Rivera Lam, M. P., Steering AI and advanced ICTs for knowledge societies: A rights, openness, access, and multi-stakeholder perspective (UNESCO series on internet freedom). (UNESCO, 2019).
- 14.Friemel, T. N. The digital divide has grown old: Determinants of a digital divide among seniors. New. Media Soc.18 (2), 313–331 (2014). [Google Scholar]
- 15.Fisk, R. P. et al. Healing the Digital Divide With Digital Inclusion: Enabling Human Capabilities. J. Service Res.26 (4), 542–559 (2022). [Google Scholar]
- 16.O’Mara, B. Social media, digital video and health promotion in a culturally and linguistically diverse Australia. Health Promot. Int.28 (3), 466–476 (2013). [DOI] [PubMed] [Google Scholar]
- 17.Alam, K. & Imran, S. The digital divide and social inclusion among refugee migrants: A case in regional Australia. Inform. Technol. People. 28 (2), 344–365 (2015). [Google Scholar]
- 18.Bender, E. M., Gebru, T., McMillan-Major, A. & Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? 列 In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21) 610–623. (Association for Computing Machinery, 2021).
- 19.Beaunoyer, E., Dupéré, S. & Guitton, M. J. COVID-19 and digital inequalities: Reciprocal impacts and mitigation strategies. Comput. Hum. Behav.111, 106424 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Ayre, J. et al. Main COVID-19 information sources in a culturally and linguistically diverse community in Sydney, Australia: A cross-sectional survey. Patient Educ. Couns.105 (8), 2793–2800 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Baganz, E., McMahon, T., Khorana, S., Magee, L. & Culos, I. Life would have been harder, harder and more in chaos, if there wasn’t internet’: digital inclusion among newly arrived refugees in Australia during the COVID-19 pandemic. Communication Res. Pract.11 (1), 24–43 (2024). [Google Scholar]
- 22.Noble, S. U. Algorithms of Oppression: How Search Engines Reinforce Racism (NYU, 2018). [DOI] [PubMed]
- 23.Liu, C. C., Gurevych, I. & Korhonen, A. Culturally aware and adapted NLP: A taxonomy and a survey of the state of the art (TACL pre-MIT Press publication version; Version 2). (2024).
- 24.Earley, P. C. & Ang, S. Cultural Intelligence: Individual Interactions across Cultures (Stanford University Press, 2003).
- 25.Van Dyne, L. et al. Sub-dimensions of the four factor model of cultural intelligence: Expanding the conceptualization and measurement of cultural intelligence. Cross Cult. Manag. Int. J. (2012).
- 26.Livermore, D., Van Dyne, L. & Ang, S. Organizational CQ: Cultural intelligence for 21st-century organizations. Bus. Horiz.65 (5), 671–680 (2022). [Google Scholar]
- 27.Capone, A. Intercultural Pragmatics. Australian J. Linguistics. 34 (2), 293–301 (2014). [Google Scholar]
- 28.McConachy, T. & Spencer-Oatey, H. Cross-cultural and intercultural pragmatics. In (eds Haugh, M., Kádár, D. Z. & Terkourafi, M.) The Cambridge Handbook of Sociopragmatics (733–757). (Cambridge University Press, 2021).
- 29.Ting-Toomey, S. Understanding Intercultural Communication (Oxford University Press, 2021).
- 30.Tenzer, H., Pudelko, M. & Harzing, A. W. The impact of language barriers on trust formation in multinational teams. J. Int. Bus. Stud.45, 508–535 (2013). [Google Scholar]
- 31.Pawar, S. et al. Survey of cultural awareness in language models: Text and beyond. https://arXiv.org/abs/2411.00860. (2024).
- 32.Iskhakova, M. Does cross-cultural competence matter when going global? Cultural intelligence and its impact on performance of international students in Australia. J. Intercultural Communication Res.47 (2), 121–140 (2018). [Google Scholar]
- 33.Cohen, L. & Kassis-Henderson, J. Revisiting culture and language in global management teams: Toward a multilingual turn. Int. J. Cross Cult. Management: CCM. 17 (1), 7–22 (2017). [Google Scholar]
- 34.Le, H., Jiang, Z. & Nielsen, I. Cognitive Cultural Intelligence and Life Satisfaction of Migrant Workers: The Roles of Career Engagement and Social Injustice. Soc. Indic. Res.139 (1), 237–257 (2018). [Google Scholar]
- 35.Kubicek, A., Bhanugopan, R. & O’Neill, G. How does cultural intelligence affect organisational culture: the mediating role of cross-cultural role conflict, ambiguity, and overload. Int. J. Hum. Resource Manage.30 (7), 1059–1083 (2019). [Google Scholar]
- 36.Australian Government Department of Home Affairs. Australia’s Multicultural Framework Review (Final Report, 2021).
- 37.OECD. Effectively Managing Investments in Digital Government: An OECD Policy Framework, OECD Public Governance Policy Papers, No. 76 (OECD Publishing, 2025).
- 38.Peters, U. & Carman, M. Cultural bias in explainable AI research: A systematic analysis. J. Artif. Intell. Res.79, 14888 (2024). [Google Scholar]
- 39.Perera, M. et al. Indigenous peoples and artificial intelligence: A systematic review and future directions. Big Data Soc.12(2). (2025).
- 40.Khatri, C. et al. Alexa prize-state of the art in conversational AI. AI Magazine. 39 (3), 40–55 (2018). [Google Scholar]
- 41.Mariani, M. M., Hashemi, N. & Wirtz, J. Artificial intelligence empowered conversational agents: A systematic literature review and research agenda. J. Bus. Res.161, 113838 (2023). [Google Scholar]
- 42.Dai, D. W., Suzuki, S. & Chen, G. Generative AI for professional communication training in intercultural contexts: Where are we now and where are we heading? Appl. Linguistics Rev.16 (2), 763–774 (2025). [Google Scholar]
- 43.Noy, S. & Zhang, W. Experimental evidence on the productivity effects of generative artificial intelligence. Science381 (6654), 187–192 (2023). [DOI] [PubMed] [Google Scholar]
- 44.Diederich, S., Brendel, A. B., Morana, S. & Kolbe, L. On the design of and interaction with conversational agents: An organizing and assessing review of human–computer interaction research. J. Association Inform. Syst.23 (1), 96–138 (2022). [Google Scholar]
- 45.Zheng, Q., Tang, Y., Liu, Y., Liu, W. & Huang, Y. UX research on conversational human–AI interaction: A literature review of the ACM Digital Library. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22)570, 1–24. (Association for Computing Machinery, 2022).
- 46.Hohenstein, J. et al. Artificial intelligence in communication impacts language and social relationships. Sci. Rep.13, 5487 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Qin, L. et al. A survey of multilingual large language models. Patterns6 (1), 101118 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Sap, M. et al. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics 5477–5490. (Association for Computational Linguistics, 2020).
- 49.Ribino, P. The role of politeness in human–machine interactions: A systematic literature review and future perspectives. Artif. Intell. Rev.56 (Suppl. 1), 445–482 (2023). [Google Scholar]
- 50.Cai, N., Gao, S. & Yan, J. How the communication style of chatbots influences consumers’ satisfaction, trust, and engagement in the context of service failure. Humanit. Social Sci. Commun.11, 687 (2024). [Google Scholar]
- 51.Poivet, R., Lopez Malet, M., Pelachaud, C. & Auvray, M. The influence of conversational agents’ role and communication style on user experience. Front. Psychol.14, 1266186. (2023). [DOI] [PMC free article] [PubMed]
- 52.Deriu, J. et al. Survey on Evaluation Methods for Dialogue Systems (Artificial Intelligence Review, 2020). [DOI] [PMC free article] [PubMed]
- 53.Rastogi, A., Zang, X., Sunkara, S., Gupta, R. & Khaitan, P. Towards scalable multi-domain conversational agents: The Schema-Guided Dialogue dataset. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-20). (2020).
- 54.Tudor Car, L. et al. Conversational Agents in Health Care: Scoping Review and Conceptual Analysis. J. Med. Internet. Res., 22(8), e17158. (2020). [DOI] [PMC free article] [PubMed]
- 55.Fang, X., Che, S., Mao, M., Zhang, J. & Wang, Y. Bias of AI-generated content: An examination of news produced by large language models. Sci. Rep.14, 5224 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Conia, S. et al. Towards cross-cultural machine translation with retrieval-augmented generation from multilingual knowledge graphs. https://arXiv.org/abs/2410.14057. (2024).
- 57.Chen, B., Zhang, Z., Langrené, N. & Zhu, S. Unleashing the potential of prompt engineering for large language models. Patterns6 (6), 101260 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.ABS. Census of Population and Housing: Cultural diversity (Australian Bureau of Statistics, 2022).
- 59.Lehdonvirta, V., Oksanen, A., Räsänen, P. & Blank, G. Social media, web, and panel surveys: Using non-probability samples in social and policy research. Policy Internet. 13 (1), 134–155 (2020). [Google Scholar]
- 60.Haugh, M. Impoliteness and taking offence in initial interactions. J. Pragmat.86, 36–53 (2015). [Google Scholar]
- 61.Wierzbicka, A. Imprisoned in English (Oxford University Press, 2019).
- 62.Matthews, S. & Yip, V. Cantonese: A Comprehensive Grammar (Routledge, 2011).
- 63.Chor, W. Sentence final particles as epistemic modulators in Cantonese conversations: A discourse-pragmatic perspective. J. Pragmat.129, 34–47 (2018). [CrossRef]. [Google Scholar]
- 64.Ponti, E. M. et al. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) 2362–2376. (Association for Computational Linguistics, 2020).
- 65.Tiedemann, J. Parallel data, tools and interfaces in OPUS. Proceedings of LREC 2012 (International Conference on Language Resources and Evaluation). (European Language Resources Association (ELRA), 2012).
- 66.Zhu, W. et al. Multilingual machine translation with large language models: Empirical results and analysis. In Findings of the Association for Computational Linguistics: NAACL 2765–2781. (Association for Computational Linguistics, 2024).
- 67.Yan, J. et al. GPT-4 vs. human translators: A comprehensive evaluation of translation quality across languages, domains, and expertise levels. https://arXiv.org/abs/2407.03658. (2024).
- 68.Brown, P. & Levinson, S. C. Politeness: Some Universals in Language Usage (Cambridge University Press, 1987).
- 69.Papineni, K., Roukos, S., Ward, T. & Zhu, W. J. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics 311–318. (ACL, 2002).
- 70.Ribeiro, M. T., Wu, T., Singh, S. & Guestrin, C. Language models don’t always say what they think: cultural and social factors influence LLM outputs. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023).
- 71.Matsumoto, D. Culture, context, and behavior. J. Pers.75 (6), 1285–1320 (2007). [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.






















