Measuring Bias In Large Language Models: A Comparative Evaluation of LLaMA3.2-1B and LLaMA3.1-8B Across Indian Socio-Culture Dimension
Abstract
In this paper we present a 30, 000-variant India-context bias audit to compare LLaMA3.2-1B and LLaMA3.1-8B across categories of Gender, Religion, Profession and Region that proposes the India Context Sensitivity Index (ICSI) as a category-weighted fairness metric. The larger model shows an improvement of the aggregate bias score that is statistically significant (Mann-Whitney p < 0.001) a biased-response rate that is significantly lower (χ² = 51.92, p < 0.001, 94.04% unbiased responses) at approximately double the inference latency. The breakdown of results by categories also shows that this overall gain is unevenly split: the 8B remains more sensitive on Gender-based prompts even as it improves on Religion, Profession and Region, underscoring the merit of reporting disaggregated, category-imbued fairness over a single number bias score for India deployed LLMs [9], [17], [23]. The future work will allow the scaling of the benchmark to be done to additional model sizes within the LLaMA family (for instance 3B, 70B) to instead fit a bias-versus-scale curve rather than a two-point comparison (14). Then, the extension of the category set to ‘caste-adjacent’ and intersectional categories (for instance gender × region) that are under-represented in this effort (18), (9). The next direction will be replacing the lexicon-based bias scorer with the LLM-as-judge scorer validated against human annotation, to measure scorer-induced bias in the evaluation pipeline itself (28). Finally, deploying the mitigation variants (few-shot, prompt-engineering) evaluated here as live runtime guardrails and measuring their effect on the ICSI in a closed deployment loop (29), (30).