Cross-Cultural Scenario Benchmark: Evaluating LLMs’ Cross-Cultural Understanding
Abstract
Cross-cultural reasoning and alignment have been identified as key weaknesses of large language models (LLMs), but the architectural or cognitive features underlying these failures have not been adequately examined. In addition, previous studies rely almost exclusively on datasets and benchmarks constructed under the WEIRD (Western, Educated, Industrialized, Rich, and Democratic) bias. To address this data bias issue, we prepare a dataset with substantial coverage of non-WEIRD cultures and five-dimensional (W, E, I, R, and D) annotations. This dataset supports an interpretable approach to examining weaknesses in LLMs’ cross-cultural alignment. We adopt the Chinese–Foreign Cultural Differences Case Repository at Xiamen University, which contains 9342 cases across 151 countries, 6 continents, and 10 cultural domains. These cases are processed and transformed into benchmark-ready structured data through topic normalization, structured metadata cleaning, continent correction, and country-level WEIRD annotation along five dimensions. Each case is converted into a six-option cultural attribution question with five cognitive-trap distractors grounded in cognitive reasoning and pragmatic interpretation. Evaluation of 6 mainstream large language models shows that their dominant failures do not involve explicit stereotypes. Instead, 61% of all errors arise from oversimplifying complex cultural phenomena or applying familiar cultural frames. The proportion of errors that explain specific cultural conflicts through seemingly universal value frames increases from 11% at the low-WEIRD end to 20% at the high-WEIRD end of the dataset. These results suggest that WEIRD data bias reflects both the underrepresentation of low-WEIRD cultures and the overactivation of dominant value frames in high-WEIRD contexts.