Skip to content

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation

Jul 2026 · arXiv.org · Vol abs/2607.20174 · 2 citations · 52 references
Computer Science

TL;DR

StreamHOI, a low-latency streaming framework for long-duration HOI video generation, finds that the standard sink-local memory design faces a trade-off in streaming HOI generation, and different transformer blocks show different historical-memory preferences for HOI regions and surrounding regions.

Abstract

Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex driving conditions, making them unsuitable for real-time interactive applications. We present \emph{StreamHOI}, a low-latency streaming framework for long-duration HOI video generation. Instead of converting heavily conditioned HOI pipelines into streaming systems, we study how an image-to-video streaming generator should organize historical memory to preserve interactions under bounded latency. We find that the standard sink-local memory design faces a trade-off in streaming HOI generation, and different transformer blocks show different historical-memory preferences for HOI regions and surrounding regions. To match memory composition with block behavior, StreamHOI performs offline HOI-aware block profiling and applies bias-guided memory-specialized training to adapt the generator to block-specific memory layouts. We further introduce a memory distance scaling module to strengthen long-range access to early interaction states. Extensive comparisons with both long-video baselines and recent HOI generation methods demonstrate that StreamHOI achieves strong interaction plausibility, object fidelity, human quality and efficiency, reaching 17.6 FPS with 0.75s first-chunk latency.

View source

Similar papers

Preprint Aug 2026

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

StreamFlow is introduced, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information and improves the visual attention score while reducing end-to-end latency and peak memory, enabling more visually grounded and efficient reasoning.

Muxin Fu, Yifan Zhang, Wentao Zhang et al. · 0 citations
Preprint Aug 2026

StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation

Recent advancements in generative models are pushing video generation toward unbounded streaming audio-video generation for real-time interactive worlds. However, existing benchmarks primarily evaluate completed sequences and struggle to capture streaming properties. To bridge this gap, we introduce StreamAV-Bench, the...

Kaiqi Liu, Hao-Xuan Zeng, Jingqi Liu et al. · 3 citations · ⚡2
Preprint Sep 2026

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require mem...

Jian-Guo Huang, Jin-Ming Liu, Qi-Yao Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserv...

Hao-Cheng Xi, Yiming Xie, He-Xue Zhao et al. · 1 citation
Preprint Aug 2026

Encore: Infinite Audio-Video Generation with Adaptive Signal Routing

Existing audio-video generation methods produce well-synchronized clips but are limited to short durations, while long-video generation methods extend duration through chunk-based iterative synthesis yet lack audio entirely. Generating long audio-video jointly is fundamentally harder than either task alone: each chunk...

Shao-Hua Pan, Junbao Chen, Shengyi He et al. · 0 citations
Preprint Aug 2026

StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models

Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational reduction. Most methods focus on optimizing the injection procedure of current data (write) and retrieving informative historical data (read) from memory, while overlooki...

Yu-Xing Liu, Peiqin Zhuang, Ya-Li Wang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.