This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state, and introduces a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse categories.
Abstract
Large language models have advanced long-form story generation, yet the resulting narratives often fail to preserve facts, states, and relationships established earlier in the same text. This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state. Using a targeted evidence-mapping strategy, we analyze 90 papers, with a literature search cutoff of April 23, 2026, covering long-form story generation, narrative evaluation, consistency detection, and mitigation strategies. We distinguish narrative consistency from factuality, faithfulness, surface coherence, and hallucination by centering the dynamically accumulated evidence of the story itself. The survey organizes consistency judgments around three evidence sources, namely story-internal propositions, source-or-canon evidence, and external-world knowledge, as determined by task conditions that specify which source is binding. We introduce a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse categories. We use the taxonomy as a common reference frame for analyzing benchmark coverage, detection methods, mitigation strategies, and open challenges, highlighting where current work concentrates and where coverage remains thin. We also separate legitimate creative extension from task-inconsistent additions that contradict, revise, or exceed the permitted generation setting. The survey closes by identifying needs for evidence-grounded oracles, subtype-aware calibration, omission and discourse-level failure evaluation, and task-conditioned verification.
Long-form fiction writers need memory that answers multi-hop questions about evolving story state: who knows a secret and when they learned it, whether an event preceded the narration that revealed it, whether a setup paid off, and how a relationship shifted. General-purpose retrieval and agent-memory systems represent entities and facts but not the narratological structure these questions turn on, so they surface the wrong evidence or none at all. We introduce the Narrative World Model (NWM), a writer-memory system that pairs a narratology-grounded typed temporal-state graph with query-conditioned hybrid retrieval. To measure memory rather than the answerer, we read every system through a single held-constant Opus 4.8 reader over only that system's chapter-safe evidence, on a reproducible public corpus and a validated multi-hop benchmark, and we compare against the strongest existing temporal-knowledge-graph agent-memory framework, Graphiti/Zep (Rasmussen et al., 2025). NWM substantially and significantly outperforms this baseline on multi-hop narratological QA across both corpora, and far exceeds GraphRAG and flat retrieval. The advantage is representational rather than an artifact of extraction: it survives rebuilding the baseline with NWM's own extractor, and traces to its narratology-grounded structure and query-conditioned retrieval, not to graph size or extractor quality.
M. Saifullah, Thomas Kornmaier, Taaha Kazi et al.· 1 citation· ⚡1
This system for the Narrative Similarity task at SemEval-2026 (Task 4), where the goal is to determine which of two candidate stories is more similar to an anchor story directly or via vector representations, finds that chain-of-thought–style prompting with detailed reasoning outputs achieves comparable results to the scoring approach on difficult examples.
Tisa Islam Erana, Azwad Anjum Islam, Anshu Kiran Sharma et al.· SemEval@ACL· 1 citation· ⚡1
NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.
Yuheng Huang, Jianlang Chen, Jiayang Song et al.· 0 citations
This work proposes StructAgent, a document-oriented LLM agent that integrates structure-aware navigation, sequential reading, and evidence-constrained generation into a unified agentic loop, and generates answers whose reasoning chains are explicitly grounded in cited evidence.
Jingfeng Zhou, Zhizhen Zhu, A. Zhu et al.· 2026 6th International Confe...· 1 citation
CraftAlign is introduced, a framework that aligns AI stories with the craft of human storytelling by both assessing Human/AI writing patterns and providing revision guidance by both assessing Human/AI writing patterns and providing revision guidance.
Yang Yang, Boyun Xu, Shaofeng Liang et al.· 0 citations
It is suggested that LLM behavior that survives narrative changes should be grounded in concrete actions rather than abstract descriptions, and that persona effects that do transfer across narratives arise from behavioral anchors, persona descriptions whose language maps directly onto shared actions.
Yixuan Wang, James C. Lester, Shashank Srivastava· 0 citations