Skip to content
Open access

Long-Horizon Video Generation with Temporally Consistent Diffusion and Scene-Graph Guidance

Aug 2026 · Journal of innovative research and technology · 0 citations

TL;DR

A novel framework that integrates temporally consistent diffusion models with dynamic scene-graph guidance that structurally constrains the generative process, ensuring that objects, their attributes, and their interrelationships remain stable over extended durations is introduced.

Abstract

The synthesis of high-fidelity, temporally coherent long-horizon videos remains a profound challenge in the domain of generative artificial intelligence. Current diffusion-based approaches often suffer from severe temporal degradation, semantic drift, and structural inconsistency when generating sequences beyond a few seconds. To address these limitations, this paper introduces a novel framework that integrates temporally consistent diffusion models with dynamic scene-graph guidance. By leveraging scene graphs as explicit semantic anchors across frames, the proposed architecture structurally constrains the generative process, ensuring that objects, their attributes, and their interrelationships remain stable over extended durations. The methodology involves a dual-stream architecture where a graph neural network processes sequential scene graphs to condition a cascaded video diffusion model. Furthermore, a specialized spatiotemporal cross-attention mechanism is developed to align latent noise representations with the relational data embedded in the scene graphs. Extensive empirical evaluations on standard video generation benchmarks demonstrate that the proposed method significantly outperforms baseline approaches in both quantitative metrics and qualitative human assessments, particularly in maintaining entity persistence and logical action progression over long time horizons. The findings underscore the critical role of explicit structural representations in overcoming the inherent memory limitations of pure attention-based video generation systems

Read PDF

Similar papers

Conference Open access 2026

Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge

Recent proprietary video generation models have demonstrated remarkable proficiency in synthesizing highly realistic videos from textual instructions. Most open-source text-to-video models, however, still struggle to accurately simulate real-world physics and dynamic entity interactions. Existing approaches rely on scaling laws and large-scale, high-quality video datasets to implicitly learn physical dynamics, yet this paradigm is constrained by prohibitive costs and the burdensome demands of data curation. Motivated by this, we propose a novel framework that integrates graph-structured temporal knowledge into video latent diffusion models to enhance compositional generation and interaction fidelity. Our framework constructs video scene graphs specifically designed to capture entity relationships, temporal dynamics, and global scene context. These graph-structured representations guide the generation process through cross-attention mechanisms. Additionally, we introduce Graph-Aligned De-noising Loss (GADL), a training objective that ensures adherence to conditioned graphs by incorporating node modification tasks within the denoising process, leveraging synchronized edited video-graph pairs. Comprehensive evaluations demonstrate that incorporating graph-structured knowledge significantly enhances compositionality and the accurate portrayal of real-world interactions in generated videos.

Joonhyung Park, Jae-gyun Song, Si-hun Park et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT is introduced, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation, and a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference.

Yushe Cao, Shikun Feng, Ru-Xiang Duan et al. · 0 citations
Preprint Jul 2026

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing

Video-text temporal localization requires precise alignment between natural language queries and corresponding video segments, a fundamental challenge in multimodal understanding. We present a novel framework that addresses two critical limitations of existing methods: inadequate modeling of hierarchical temporal structure and inability to handle complex many-to-many correspondences between modalities. Our approach introduces a multi-scale temporal convolutional encoder that captures motion patterns across different temporal granularities - from instantaneous frame transitions to extended action sequences. We further propose a capsule-based dynamic routing mechanism that iteratively refines segment-query associations through structured agreement updates, enabling flexible modeling of non-monotonic alignments. These components are unified through a multi-task learning objective that jointly optimizes temporal boundary regression, cross-modal semantic alignment, and capsule diversity. Extensive experiments on ActivityNet Captions demonstrate significant improvements, achieving 42.9% Recall@0.5 and 41.1% mean IoU, surpassing strong transformer-based baselines while maintaining computational efficiency. Our results validate that combining hierarchical temporal modeling with structured semantic routing provides an effective solution for fine-grained video-language understanding.

Gengtian Shi, Jinze Yu, Chenhao Wu et al. · 0 citations
Preprint Aug 2026

Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

This work proposes an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes and consistently outperforms existing state-of-the-art approaches.

Guixu Lin, Yuyang Yu, Xiang Ji et al. · 0 citations
Preprint Jul 2026

Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency

This work proposes Cycle-World, a novel framework designed for stable and temporally consistent long-video generation that tackles error drift by enforcing strict temporal reversibility across both the training and inference phases, and demonstrates that forward generative drift can be strictly bottlenecked by a cycle-consistency objective.

Zihan Su, Teng Hu, Jiangning Zhang et al. · 1 citation