Multi-Test Nation Bias Benchmarking of LLMs with Explicit Unbiased Baselines
While prior work has highlighted that systematic nation-level bias in large language models (LLMs) can pose operational risks for international relations (IR) applications, many existing evaluations still lack a clearly specified unbiased reference (ground truth), limiting fully quantitative and cross-setting measure...