Modeling social context in natural language processing
Hundreds of millions of people interact with language models (LMs) every day, using them for tasks such as writing assistance and information seeking. As their range of applications grows, it becomes increasingly essential to consider the role of social context in how these systems are designed and deployed. In particular, this thesis focuses on LMs' sociolinguistic competence—how associations between linguistic variation and social dimensions can be incorporated into LMs, how they are learned and manifested, and how they can lead to harm. In the first part of the thesis, we develop computational methods that improve LMs' sociolinguistic competence by explicitly injecting social context into the model architecture. We focus on two forms of social context—social networks and geographic location—and draw on recent advances in graph neural networks and multi-task learning to integrate them into LMs. Across a range of benchmarks, the proposed methods yield substantial gains. In the second part of the thesis, we explore whether LMs acquire sociolinguistic competence as a by-product of pretraining and posttraining, without being explicitly conditioned on social context. Experiments on dialectal variation and ideological framing suggest that LMs indeed learn associations between linguistic variation and social dimensions, albeit with varying levels of detail. Beyond these sociolinguistic associations, we also examine the question of how LMs' outputs reflect ideological leanings more generally, finding substantial evidence of instability. In the third part of the thesis, we investigate the harms that associations between linguistic variation and social dimensions can produce in LMs. Focusing on African American English, we find that LMs associate its speakers with pernicious stereotypes triggered by linguistic features alone, and that current posttraining practices do not address this covert racism. Preventing such harms is a critical goal for future research to ensure safe and equitable language technology. Finally, we release new analysis tools and datasets that facilitate broader empirical study of social context in natural language processing and computational social science, supporting subsequent work in these areas.