Skip to content
Preprint

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

Jul 2026 · 0 citations · 52 references
Computer Science

TL;DR

UniCaMo is presented, a unified framework that enables simultaneous control of object trajectories and camera viewpoints by directly constructing the input noise of the diffusion model and achieves state-of-the-art results in both video quality and motion controllability on standard controllable video generation benchmarks.

Abstract

Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approaches still lack precise and unified controllability over both object motion and camera motion within a single generation process. We present UniCaMo, a unified framework that enables simultaneous control of object trajectories and camera viewpoints by directly constructing the input noise of the diffusion model. Specifically, UniCaMo builds a shared 3D-grounded motion-consistent noise space across latent video frames. Sparse 3D point tracks are used to warp the Gaussian noise of the reference frame along desired object trajectories, while a virtual spherical noise representation provides globally consistent noise values for newly revealed scene regions under camera motion. By combining local track-guided noise warping with global sphere-based noise sampling, UniCaMo maintains geometric and temporal consistency under both object movement and viewpoint changes. Because UniCaMo modifies only the input noise, it requires no auxiliary adapters, control branches, or architectural changes to the underlying video diffusion model. With lightweight LoRA fine-tuning on large pretrained video diffusion models, including Wan 2.1 (14B), UniCaMo achieves state-of-the-art results in both video quality and motion controllability on standard controllable video generation benchmarks.

View source

Similar papers

Jul 2026

UniCam: Taming Unified Diffusion Models in Noise Space for Camera-controllable Video Rendering

The UniCam framework is proposed, a unified framework that introduces a temporally coherent stochastic representation, termed CameraNoise, warped from camera intrinsic and extrinsic parameters, which significantly outperforms prior methods in both fidelity and controllability.

Haoyu Zhao, Zuxuan Wu, Yu-Gang Jiang · 0 citations
Preprint Jul 2026

World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

The method sets a new state of the art in 4D reconstruction and seamlessly generalizes to in-the-wild videos with large viewpoint changes and dynamic motions, improving both novel-view synthesis and the underlying 3D motion.

Liyuan Zhu, Shengyu Huang, Amrita Mazumdar et al. · 0 citations
Preprint Jul 2026

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-position correspondence as a practical 4D rendering condition.

Junhao Chen, Mingjin Chen, He Zhang et al. · 0 citations
Preprint Jul 2026

OmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields

OmniX achieves state-of-the-art performance on dense 3D point trajectory prediction and 3D point tracking, while also demonstrating competitive results on video depth estimation and camera pose estimation.

Yanqin Jiang, Tengfei Wang, Zhengwei Wang et al. · 2 citations
Jun 2026

3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

A ground-adaptive 3D motion retargeting approach to enable user-friendly motion trajectory control adapting to the changes of elevations of ground and orientations automatically and a viewpoint-adaptive latent fusion mechanism to inject point-cloud geometric priors through scene-visibility masking into the generative process, providing precise guidance of viewpoint changes under camera control are developed.

Deyin Liu, Jichen Xu, L. Wu et al. · 0 citations
Jun 2026

Again-Pose: Anchor-Guided Adaptive Inter-Frame Motion Cues Propagating for High-quality Human Pose Reconstruction

This work proposes a simple yet effective framework called Anchor-guided adaptive inter-frame motion cues propagating (Again-Pose), reformulating pose estimation in degraded frames as a motion-guided recovery task, significantly outperforms state-of-the-art methods in robustness and stability.

Shuaikang Zhu, Yiding Sun, Yang Yang · 0 citations