Evaluating Cultural Balance and Representation in ChatGPT and DeepSeek-Generated EFL Materials
Generative artificial intelligence is increasingly used to create English as a Foreign Language (EFL) reading materials, yet fluent output may still simplify cultural groups. This study compared 24 classroom-oriented passages generated by ChatGPT GPT-5.6 and DeepSeek-V4 from 12 identical prompts submitted on 22 July 2026. Each prompt requested a 220–250-word intermediate-level passage representing Chinese, Pakistani, and mainstream English-speaking Western perspectives fairly. Directed qualitative content analysis examined inclusion, balance, intragroup variation, specificity, essentialization, the collectivist-individualist binary, evaluative hierarchy, stereotyping risk, intercultural sensitivity, and prompt compliance. ChatGPT produced 2,763 words and met the requested range in all 12 passages. DeepSeek produced 3,095 words and met the requested range in five passages. Both systems included the three perspectives and promoted respectful communication. ChatGPT used more qualifications and acknowledged internal variation more consistently. DeepSeek supplied more named cultural detail but relied more often on broad contrasts that framed Chinese and Pakistani contexts as collective and Western contexts as individualistic. The findings show that cultural inclusion does not ensure balanced representation. The proposed review criteria help teachers and curriculum developers evaluate variation, specificity, stereotyping risk, and classroom suitability before using AI-generated materials.