Skip to content
Review Open access

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Aug 2026 · Digital Studies in Language and Literature · 0 citations · 40 references

TL;DR

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Abstract

Abstract Large language models (LLMs) are increasingly used not only to retrieve information but also to generate language on behalf of users, raising important questions about how they mediate cultural identity and authorial voice. This study investigates that issue through a comparative analysis of five widely used systems – ChatGPT, Claude, Copilot, Gemini, and Perplexity – using the Inglehart–Welzel cultural map as an empirical framework. The analysis is conducted in two stages: first, a quantitative assessment of cultural value alignment based on model-generated responses to standardized survey items under neutral and country-specific prompting conditions; and second, a qualitative examination of linguistic features in selected excerpts of responses to explore how voice and stance are constructed. Representational alignment is measured as the distance between model outputs and national cultural profiles. The results show that baseline responses cluster around English-speaking and European cultural regions, indicating shared default priors, while country-specific prompting improves alignment, though performance is uneven across models and persistent gaps remain in several non-Western contexts. These patterns are conceptualized as cultural shadows, reflecting the influence of dominant cultural norms in model outputs. Drawing on Hyland’s (2005) account of stance and McKinley’s (2026) view of generative artificial intelligence (GenAI) as a rhetorical force, the findings suggest that these patterns extend beyond value alignment to the mediation of voice and self-representation, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Read PDF

Similar papers

Review Open access Aug 2026

Evaluating Cultural Balance and Representation in ChatGPT and DeepSeek-Generated EFL Materials

Generative artificial intelligence is increasingly used to create English as a Foreign Language (EFL) reading materials, yet fluent output may still simplify cultural groups. This study compared 24 classroom-oriented passages generated by ChatGPT GPT-5.6 and DeepSeek-V4 from 12 identical prompts submitted on 22 July 2026. Each prompt requested a 220–250-word intermediate-level passage representing Chinese, Pakistani, and mainstream English-speaking Western perspectives fairly. Directed qualitative content analysis examined inclusion, balance, intragroup variation, specificity, essentialization, the collectivist-individualist binary, evaluative hierarchy, stereotyping risk, intercultural sensitivity, and prompt compliance. ChatGPT produced 2,763 words and met the requested range in all 12 passages. DeepSeek produced 3,095 words and met the requested range in five passages. Both systems included the three perspectives and promoted respectful communication. ChatGPT used more qualifications and acknowledged internal variation more consistently. DeepSeek supplied more named cultural detail but relied more often on broad contrasts that framed Chinese and Pakistani contexts as collective and Western contexts as individualistic. The findings show that cultural inclusion does not ensure balanced representation. The proposed review criteria help teachers and curriculum developers evaluate variation, specificity, stereotyping risk, and classroom suitability before using AI-generated materials.

Idrees, Liu Yongzhi · 0 citations
Preprint Aug 2026

Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning

Figurative language is deeply culturally embedded; fluent use requires not just linguistic competence but cultural immersion. We ask whether LLMs can learn this link: does fine-tuning on cultural data improve figurative language understanding, and vice versa? We conduct a systematic study across four models (ALLaM-7B, Fanar-1-9B, Qwen3-8B, Llama-3.1-8B) and six Arabic datasets spanning cultural commonsense, proverbs, and poetry across diverse dialects and regions. Fine-tuning on poetry improves idiom comprehension (+2.33%, p<0.05), a gain our ArabicMMLU control does not reproduce, indicating that it stems from figurative content rather than Arabic language adaptation and pointing to a sensitivity to non-literal meaning that transfers across figurative types. Cultural fine-tuning, by contrast, lowers proverb-interpretation accuracy in both Arabic-centric models. Transfer between the two domains is otherwise indistinguishable from noise, with Arabic models frequently regressing after fine-tuning, suggesting prior saturation of relevant knowledge, while multilingual models show greater adaptation headroom. Error analysis further reveals that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge. Our findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone.

Mena Attia, Mona T. Diab, T. Solorio · 0 citations
Preprint Aug 2026

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.

K. V. Koevering, Anjalie Field · 0 citations
Preprint Aug 2026

Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs

Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently. We ask where inside the model this cultural default is produced. On a parallel cross-cultural substrate of Thompson-motif entities, we instrument 18 open-source LLMs from 8 architecture families with linear probing, logit lens, activation patching, and output extraction. The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation. Asking the same question in the target culture's native language versus English produces failures that cluster within language but decouple across language: the decoder is gated on prompt language. We release a per-entity (probe, output) decomposition framework, a citation-anchored cross-cultural ground truth, a within- versus cross-mode correlation test for language-conditioned readout, and per-entity predictions for all 18 models.

Iaroslav Chelombitko, Ekaterina Chelombitko, Mika Hämäläinen · 0 citations
Open access Aug 2026

Cross-Cultural Scenario Benchmark: Evaluating LLMs’ Cross-Cultural Understanding

Cross-cultural reasoning and alignment have been identified as key weaknesses of large language models (LLMs), but the architectural or cognitive features underlying these failures have not been adequately examined. In addition, previous studies rely almost exclusively on datasets and benchmarks constructed under the WEIRD (Western, Educated, Industrialized, Rich, and Democratic) bias. To address this data bias issue, we prepare a dataset with substantial coverage of non-WEIRD cultures and five-dimensional (W, E, I, R, and D) annotations. This dataset supports an interpretable approach to examining weaknesses in LLMs’ cross-cultural alignment. We adopt the Chinese–Foreign Cultural Differences Case Repository at Xiamen University, which contains 9342 cases across 151 countries, 6 continents, and 10 cultural domains. These cases are processed and transformed into benchmark-ready structured data through topic normalization, structured metadata cleaning, continent correction, and country-level WEIRD annotation along five dimensions. Each case is converted into a six-option cultural attribution question with five cognitive-trap distractors grounded in cognitive reasoning and pragmatic interpretation. Evaluation of 6 mainstream large language models shows that their dominant failures do not involve explicit stereotypes. Instead, 61% of all errors arise from oversimplifying complex cultural phenomena or applying familiar cultural frames. The proportion of errors that explain specific cultural conflicts through seemingly universal value frames increases from 11% at the low-WEIRD end to 20% at the high-WEIRD end of the dataset. These results suggest that WEIRD data bias reflects both the underrepresentation of low-WEIRD cultures and the overactivation of dominant value frames in high-WEIRD contexts.

Meng-Xi Guo, Lei-Ming Gao, Winnie Zeng et al. · 0 citations