Skip to content
Preprint

Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation

Aug 2026 · 3 citations · 59 references
Computer Science

TL;DR

Context-Matched Distillation (CMD) is introduced, a causal DMD framework that aligns teacher supervision with the information available when each target is generated, and naturally extends to frame-wise and chunk-wise generation, long video distillation, and camera-conditioned distillation.

Abstract

Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step distillation accelerates generation by reducing denoising steps, while online control imposes a causal constraint: frames and blocks should depend on history and controls available during generation. Existing video distribution matching distillation (DMD) pipelines, however, often supervise causal few-step students using bidirectional teachers that score complete clips. The score for a target can therefore depend on future frames and controls that were unavailable when the student generated it, misaligning teacher supervision with the student's causal information set. We introduce Context-Matched Distillation (CMD), a causal DMD framework that aligns teacher supervision with the information available when each target is generated. CMD replaces bidirectional full-clip scoring with a causal teacher that evaluates each target without access to future frames or controls. The same causal teacher initializes the few-step student, establishing a consistent causal formulation across teacher training, student distillation, and inference. Beyond aligning the temporal information boundary, Prefix Scoring matches supervision to the student's realized rollout context by evaluating each target under the cached student-generated prefix that produced it. Prefix Corruption further stabilizes training by perturbing unreliable prefixes produced early in training while preserving this target-context alignment. With a simple causal formulation, CMD naturally extends to frame-wise and chunk-wise generation, long video distillation, and camera-conditioned distillation. Experiments demonstrate state-of-the-art aggregate performance among autoregressive methods on both short- and long-video benchmarks, together with substantially improved adherence to time-varying camera controls.

View source

Similar papers

Preprint Aug 2026

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must re...

Xin-Ye Li, Lingshuai Lin, Lei Wang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework mi...

Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al. · 1 citation
Preprint Aug 2026

Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts

Visual on-policy distillation (OPD) improves the training of compact visual autoregressive models by learning from trajectories generated by the current student. However, these online rollouts are still produced token by token with autoregressive decoding, which adds substantial cost to every on-policy training step. S...

Bingqi Shan, Zhehao Yu, Keng-Hong Lin et al. · 0 citations
Preprint Sep 2026

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both a large pretrained teacher and an online critic to estimate distributional discrepancies through diffusion scores. In t...

Zi-Chong Meng, Chong-Jian Ge, C. Huang et al. · 0 citations
Preprint Sep 2026

CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction

Autoregressive vision language models unify heterogeneous perception tasks but are highly susceptible to compounding errors. On-policy distillation (OPD) bridges the training-inference mismatch by training students on their own rollouts. However, unreliable student predictions, especially early in training, can derail...

Meng-Hao Li, Lin-Jie Mu, Yin Wang et al. · 1 citation
Preprint Sep 2026

UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation

Distilling bidirectional multi-step video diffusion transformers into few-step causal models has become a common approach for streaming video generation. While these few-step students are significantly faster than the teachers they are distilled from, they remain slow for real-time generation. In this work we present U...

Youssef Mansour, Enis Simsar, F. Sener et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.