Skip to content
Preprint

From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives

Jul 2026 · 0 citations · 42 references
Computer Science

TL;DR

It is suggested that long-form narratives can emerge from explicit world-state tracking and goal-driven multi-agent generation, providing a foundation for controllable and structurally coherent long-form narrative generation.

Abstract

Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce a unified framework for long-form narrative generation and verification. MAGNET, a multi-agent goal-driven narrative engine for storytelling, generates stories with persona-grounded character agents that propose actions based on a shared world state and evolving story goals, while ATLAS is a graph-based pipeline that compares scene-level world representations across a generated story to detect hallucinations. By evaluating MAGNET using an LLM editor, pairwise rubric scoring, and ATLAS, we show that our framework produces coherent narratives compared to single-model prompting and IBSEN. At 100 pages, MAGNET reduced annotations and hallucinations by 41 and 50%, respectively, compared to the single model baseline and by 34 and 45%, respectively, compared to IBSEN, with pairwise rubric evaluation showing similar results. These results suggest that long-form narratives can emerge from explicit world-state tracking and goal-driven multi-agent generation, providing a foundation for controllable and structurally coherent long-form narrative generation.

View source

Similar papers

Review Open access 2026

Narrative Consistency in Large Language Model-Generated Stories: A Survey

This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state, and introduces a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse categories.

Keunhyeung Park, Seunguk Yu, Jinhee Jang et al. · 0 citations
Conference Jul 2026

Agentbook: Enabling Fine-Grained Editability in Visual Narratives via Page-Centric Agent Coordination

Long-horizon visual storytelling with text-to-image diffusion models enables coherent multi-page narratives from a single prompt, yet existing single-pass generation pipelines tightly couple pages within a shared latent trajectory, limiting structured editability. We propose AgentBook, a training-free multi-agent framework that reformulates story generation as a page-centric process governed by an explicit global narrative state encoding character identity, stylistic constraints, and narrative context. By decoupling page synthesis from a monolithic diffusion chain and coordinating story planning, generation, consistency enforcement, and user editing through state-based conditioning, AgentBook enables localized page regeneration and iterative refinement without compromising cross-page coherence. Experimental results demonstrate improvements in story coherence, visual consistency, and edit locality over representative visual storytelling and text-to-image diffusion baselines, establishing a controllable and interactive paradigm for long-form visual narrative generation.

Ayushman Sarkar, Zhenyu Yu · 0 citations
Preprint Jul 2026

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

FilmEval is introduced, a systematic evaluation framework that couples a difficulty-graded benchmark of 15 representative novels with an automated protocol of nine objective metrics spanning three dimensions: cinematic presentation, film consistency, and novel fidelity.

Jialong Zuo, Haotong Zuo, Shiwei Zhang et al. · 0 citations
Preprint Aug 2026

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.

Yuheng Huang, Jianlang Chen, Jiayang Song et al. · 0 citations