Aug 2026· Computer graphics forum (Print)· 0 citations· 51 references
TL;DR
F3AMD (Fast FiLM‐conditioned Fourier Autoregressive Motion Diffusion), a framework that achieves an order of magnitude speedup over state‐of‐the‐art systems for multi‐character animation on both GPUs and CPUs while maintaining high motion quality, is introduced.
Abstract
Recent advances in generative motion synthesis have enabled realtime autoregressive generation of diverse and realistic character animations conditioned on user inputs, as demonstrated by models such as the Conditional Autoregressive Motion Diffusion Model (CAMDM). However, real‐world applications (e.g., computer games) often demand faster‐than‐realtime performance for large numbers of characters. We introduce F3AMD (Fast FiLM‐conditioned Fourier Autoregressive Motion Diffusion), a framework that achieves an order of magnitude speedup over state‐of‐the‐art systems for multi‐character animation on both GPUs and CPUs while maintaining high motion quality. Our key insight is that autoregressive motion diffusion is primarily bottlenecked by architectural and sampling inefficiencies. To address this, F3AMD employs Fourier Neural Operators (FNOs) as encoder‐decoder modules, substitutes Transformer backbones with FNO blocks, replaces condition concatenation with lightweight Feature‐wise Linear Modulation (FiLM), and adopts a variance‐exploding noise schedule with a deterministic sampler. This design enables a substantially lower‐dimensional latent space, facilitates learning in both the spectral and temporal domains, and significantly improves sample efficiency. We conduct systematic ablations of key design factors, including latent dimension, backbone type, diffusion window length, and number of denoising steps. Our recommended configuration, F3AMD‐FNO‐96, achieves 20x speedup over the baseline CAMDM model, while maintaining comparable motion quality.
This work introduces a novel motion prior based on the sparsity of high‐order temporal derivatives, serving as a kinematic proxy for impulsive force generation and achieves linear complexity, enabling efficient processing of long sequences.
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic penalizes genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already-static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. Across various autoregressive video backbones and various generation horizons, Stream4D improves 4D reconstruction quality, preserves motion more effectively, and achieves higher human-aligned preference. Project page: https://banyuanhao.github.io/Stream4D/
Yuanhao Ban, Jiaqi Feng, Hengguang Zhou et al.· 0 citations
This work presents LiveAnimate, to their knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT).
Yuxuan Zhang, H. Xiong, Yubo Huang et al.· 0 citations
SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure and consistently improves temporal quality across multiple autoregressive diffusion models.
Thanh-Nhan Vo, Trong-Thuan Nguyen, T. Le et al.· 0 citations
AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk.
AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al.· 2 citations
MobileWan becomes the first 5B-scale video diffusion model deployable on a commercial mobile device and proposes a learnable attention head pruning method based on binary per-head gates optimized end-to-end using a noise-biased sparsity objective and distillation-based finetuning.
Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv et al.· 1 citation