Skip to content

FossilWriter: Learning Hypergraph World Models with Latent Narratives for Creative Story Generation

· 0 citations · 17 references

TL;DR

FossilWriter represents the story world as a shared hypergraph and introduces a narrative layer that marks elements with developmental potential that grows during writing and serves as the single source of truth to ensure long-range consistency.

View source

Similar papers

Review Open access 2026

Narrative Consistency in Large Language Model-Generated Stories: A Survey

This survey examines the problem as narrative consistency, defined as the task-conditioned preservation of binding propositions in the operative narrative state, and introduces a four-category, fourteen-subtype taxonomy comprising World and Setting, Character-Agentive, Event-Structural, and Narration and Discourse categories.

Keunhyeung Park, Seunguk Yu, Jinhee Jang et al. · 0 citations
Preprint Jul 2026

StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. Instead of relying only on local visual cues, StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes, enabling later descriptions to remain coherent, grounded, and contextually informative. Given only raw video and a movie title, StoryTeller can optionally retrieve public movie metadata to resolve names and story context, while accepting only facts that are supported by the video through semantic filtering and VLM verification. The method requires no subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. To evaluate whether generated AD preserves narrative information, we introduce StoryAD-QA, a question-answering benchmark that tests whether a language model can answer story-context questions using only the generated descriptions. Experiments on standard AD benchmarks and diverse long-form videos show that StoryTeller consistently improves narrative coherence, factual grounding, and story comprehension over strong baselines in automatic, QA-based, and human evaluations.

Seung-Yeon Hahm, Minh T. Dinh, SouYoung Jin · 0 citations
Preprint Aug 2026

SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

This work presents SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations, and releases PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes.

Maolin Ran, Xiaoyan Lu, Jiaqi Liu et al. · 0 citations
Preprint Aug 2026

CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories

CraftAlign is introduced, a framework that aligns AI stories with the craft of human storytelling by both assessing Human/AI writing patterns and providing revision guidance by both assessing Human/AI writing patterns and providing revision guidance.

Yang Yang, Boyun Xu, Shaofeng Liang et al. · 0 citations
Preprint Aug 2026

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.

Yuqi Chen, Sixuan Li, Yunfeng Cai et al. · 0 citations
Conference Jul 2026

Agentbook: Enabling Fine-Grained Editability in Visual Narratives via Page-Centric Agent Coordination

Long-horizon visual storytelling with text-to-image diffusion models enables coherent multi-page narratives from a single prompt, yet existing single-pass generation pipelines tightly couple pages within a shared latent trajectory, limiting structured editability. We propose AgentBook, a training-free multi-agent framework that reformulates story generation as a page-centric process governed by an explicit global narrative state encoding character identity, stylistic constraints, and narrative context. By decoupling page synthesis from a monolithic diffusion chain and coordinating story planning, generation, consistency enforcement, and user editing through state-based conditioning, AgentBook enables localized page regeneration and iterative refinement without compromising cross-page coherence. Experimental results demonstrate improvements in story coherence, visual consistency, and edit locality over representative visual storytelling and text-to-image diffusion baselines, establishing a controllable and interactive paradigm for long-form visual narrative generation.

Ayushman Sarkar, Zhenyu Yu · 0 citations