Skip to content

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

Sep 2026 · 2 citations · 41 references
Computer Science

Abstract

We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trained on connected multi-prompt videos to supervise a block-level causal student through distribution-matching distillation; and (3) Low-cost real-time interaction, combining four-step generation with context-preserving streaming to support 832 x 480 inference at 24 FPS at an estimated server rental cost of approximately USD 0.009 per stream-minute. Zing-0.5 achieves an overall score of 81.0 and a consistency score of 88.5 across 158 WBench Navigation cases. A joint-control demonstration shows a text-directed event change during continued navigation without restarting generation. We release the model weights, inference code, and Zing-SGLang serving implementation to support further work on playable generated worlds.

View source

Similar papers

Preprint Aug 2026

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must re...

Xin-Ye Li, Lingshuai Lin, Lei Wang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

H3-World: Turning Language Understanding into World Control

H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model, and introduces temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions.

Dan Chen, Ze-Qing Wang, Zibin Lin et al. · 3 citations
Preprint Sep 2026

WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon

Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time responsiveness. In this pap...

Hai-Yu Zhang, Wen-Qiang Sun, Teng-Fei Wang et al. · 0 citations
#computer vision Preprint Aug 2026

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

GlanceWAM is introduced, which decouples imagination from control within a single video DiT: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate purely in latent space without blo...

Lin-Han Wang, Zi-Jian An, Mingyuan Zhang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence

In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation:...

Shilong Zou, Shi-Lin Zhang, Yingji Zhang et al. · 0 citations
Preprint Sep 2026

MixiMotion: One-Step Text-to-Motion Generation via Asymmetric Set Distillation

Iterative text-to-motion generation delivers high-quality and semantically aligned motions but requires multiple network evaluations, resulting in substantial inference latency. We present \textbf{MixiMotion}, a strict one-step text-to-motion generation framework based on offline set distillation. Instead of distilling...

H. Dinh, B. Mai, Tran Quoc Bao Le et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.