Skip to content
Preprint

DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking

Jul 2026 · 0 citations · 35 references
Computer Science

TL;DR

DeforM is proposed, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions, and introduces a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks.

Abstract

Video generation models achieve high visual quality but often struggle to generate physics-aware videos. Unlike rigid-body motion, which can be described by explicit trajectories or formulas, complex deformation dynamics remain challenging to synthesize. We observe that a lack of physical reasoning for localizing dynamic areas allows irrelevant regions to dilute the model's attention, leading to generation failure. In this paper, we propose DeforM, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions. To reason about and localize these critical regions, we introduce a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks. For physical guidance, we develop two alternative strategies: DeforM-Free for training-free mechanism analysis and DeforM-Injection as a powerful training-based generator. Experimental results demonstrate that DeforM improves the realism of generated deformation scenarios, outperforming baseline models in both visual quality and physical consistency.

View source

Similar papers

Preprint Jul 2026

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality.

Tianxi Chen, Hanmo Chen, Huajin Chen et al. · 0 citations
Preprint Jul 2026

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

PhyParam is presented, a physics-guided image-to-video diffusion model that conditions on object-level forces, masses, friction, restitution, and scene-level gravity via a lightweight physical-attention routing mechanism, and further improves motion learning with semantic-structural feature-space supervision.

Yan-Xun Li, Hao Wen, Bingze Song et al. · 0 citations
#generative ai Preprint Aug 2026

MotionPhys: Detecting AI-Generated Videos via Physical Consistency of Optical-Flow Trajectories

This work introduces MotionPhys, a lightweight and interpretable framework that treats sparse motion trajectories as physical evidence rather than relying on appearance artifacts or generator-specific traces and reveals subtle motion inconsistencies that are difficult to capture with conventional visual cues and transforms them into a compact representation for efficient detection.

Hao He, Hao Tan, Zichang Tan et al. · 0 citations
Jul 2026

Motion-driven 4D scene generation

This paper presents an innovative method that leverages user-specified action paths to guide the 4D scene generation that dynamically synchronizes motions in the action path domain with their corresponding contents in the time domain.

Guo-Wei Yang, Qun-Ce Xu, Zhao Wei et al. · 0 citations
2025

PhysDiff-VTON: Cross-Domain Physics Modeling and Trajectory Optimization for Virtual Try-On

We present PhysDiff-VTON, a diffusion-based framework for image-based virtual try-on that systematically addresses the dual challenges of garment deformation modeling and high-frequency detail preservation. The core innovation lies in integrating physics-inspired mechanisms into the diffusion process: a pose-guided deformable warping module simulates fabric dynamics by predicting spatial offsets conditioned on human pose semantics, while wavelet-enhanced feature decomposition explicitly preserves texture fidelity through frequency-aware attention. Further enhancing generation quality, a novel sampling strategy optimizes the de-noising trajectory via least action principles, enforcing temporal coherence, spatial smoothness, and multi-scale structural consistency. Comprehensive evaluations across multiple datasets demonstrate significant improvements in both geometric plausibility and perceptual quality compared to existing approaches. The framework establishes a new paradigm for synthesizing photorealistic try-on images that adhere to physical constraints while maintaining intricate garment details, advancing the practical applicability of diffusion models in fashion technology.

Shibin Mei, Bingbing Ni · 1 citation
Preprint Aug 2026

RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

RididBench is introduced, a simulator-grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description, with per-frame masks, depth, 6-DoF trajectories, and contacts available for scoring.

Swarnim Jain, Shangzhe Wu · 2 citations