Can automatic reading be close? On the implications of using a large language model to identify interactions between lighting and emotions in French literature, 1838–1928
Abstract
This article evaluates whether a generic large language model (LLM) can be used to support literary scholarship by (1) identifying relevant passages and (2) producing explanations for those selections in response to a research question. We designed a workflow that combined semantic retrieval, vector search, and GPT-4o-based evaluation, and applied it to ninety novels set in Paris between 1838 and 1928, including works by Flaubert, Zola, Colette, and Proust, as well as lesser-known authors. Asked to identify passages in which artificial light played a role in the depiction of romantic or loving feelings, the model was relatively successful. Although it recovered only about 16 per cent of examples already located through manual reading, 56 per cent of the 506 additional passages it suggested were judged potentially valid in a scholar-evaluated sample. However, the model failed to generate explanations that would be convincing close readings. Critiquing the notion of ‘automatic close reading’, we apply our own close reading to the model’s imitations of close reading, and identify four recurrent shortcomings: (1) the problem of contemporary semantics being applied to a historical corpus, (2) a possible sentimentalist bias, (3) difficulty in detecting irony, and (4) misattribution of causality. We conclude that, if LLMs were to be trained for literary research, a model would need ‘knowledge’ of the whole text, extensive cultural and historical awareness, and training in reasoning informed by literary scholarship.