Jul 2026· International Conference on Computer Graphics and Interactive Techniques· 0 citations· 49 references
Computer Science
Abstract
Motion warping is a core technique in character animation that enables the adaptation of existing motion data to novel spatio-temporal constraints. Conventional motion warping methods often rely on heuristic modifications that can violate physical consistency or introduce visual artifacts. More recent learning-based editing approaches improve realism, but many of them encode motion into tightly entangled latent space, which makes them struggle to balance editing flexibility and content preservation. To address this, we propose a novel deep motion warping framework that explicitly disentangles the motion structure from global and stylistic attributes for intuitive motion editing. Our key insight is to leverage learned phase features as a continuous and robust representation of the underlying structure, and explicitly disentangle motion into root velocity, phase, and learned latent variables using a phase-conditioned diffusion autoencoder. This design supports a wide range of editing operations, including root motion warping, motion exaggeration, time warping, and style transfer by directly manipulating decoupled components, without requiring paired training data. Extensive experiments demonstrate that our approach enables high-level, flexible motion editing while strictly preserving the structural consistency and physical plausibility of the source motion
QWERTY is introduced, a training-free framework that enables flexible motion control in pretrained image-to-video DiTs via user-defined object warping and optical flow and achieves the most effective motion control among existing training-free approaches on a recent image-to-video DiT, with performance comparable to fine-tuning-based methods.
K. Choo, Young Min Kim, Hyunkyung Han et al.· 0 citations
MotionCraft is presented, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface to deliver temporally consistent, high-quality reconstructions under streaming constraints.
Rong Fu, Chunlei Meng, Yangcheng Zeng et al.· 0 citations
This work introduces stylized phase manifolds—a compact, interpretable latent representation that disentangles motion content, the temporal structure, and style and develops a diffusion‐based motion generator that enables fine‐grained control over semantic, temporal, and stylistic aspects of motion.
Jingyuan Li, Peizhuo Li, A. Aristidou et al.· Computer graphics forum (Pri...· 0 citations
Text-to-motion generation aims to synthesize semantically consistent and naturally coherent motion sequences from natural language descriptions. Given the continuous nature of human motion, diffusion models operating in a continuous latent space offer inherent advantages over vector quantization-based methods, particularly in avoiding quantization errors and in modeling quality. However, existing diffusion models primarily rely on mean squared error loss. This stepwise regression paradigm often leads to ‘over-smoothed’ motion sequences and struggles to capture the subtle semantic nuances embedded in textual descriptions. To realize the potential for continuous diffusion generation, an enhanced latent-space diffusion framework designed to elevate generation capabilities across two dimensions, namely, distribution approximation and semantic alignment, is proposed. Specifically, a latent-space adversarial discriminator is incorporated. By applying decoupled adversarial supervision, this component mitigates the detail loss caused by mean regression, significantly enhancing the physical realism and dynamic sharpness. Concurrently, a latent-space contrastive alignment strategy is introduced during the denoising process that reinforces the correspondence of the generated motion sequences with the given textual inputs via explicit cross-modal constraints. Extensive experiments on standard benchmarks demonstrate that the proposed method effectively addresses the limitations of conventional diffusion models, thus validating the potential of continuous diffusion frameworks within the domain of text-driven motion synthesis.
Zhaowu Li, Rui Liu, Deheng Zhu et al.· Visual Computing for Industr...· 0 citations
This work proposes a novel Deconstruct-Recompose Paradigm (DRP) for learning transferable local motion representations and introduces a Dual-Attention Encoder to learn local motion representations from these Atomic Actions, capturing their spatiotemporal relationships.
Jinwen Wang, Youfang Lin, X. Hu et al.· 0 citations
This work proposes Change and Invariance Motion Editing (CIME), a unified framework that comprehensively decouples change and invariance into spatial pose and temporal rhythm dimensions and introduces the Riemannian Non-uniform Integral Manifold Mapping module.
Shaohui Lin, Zhenwu Shi, Jingyu Gong et al.· 0 citations