Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Sep 2026

The Past Frames the Future: Memory for Autoregressive Video Generation

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands,...

Harold Haodong Chen, Rong-Jin Guo, Di-Sen Lan et al. · 0 citations
Preprint Aug 2026

Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs

Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlenecks in complex visual tasks. However, existing approaches rarely verify tool outputs, limiting their ability to detect and recover from too...

Yu Cheng, A. Goel, Hakan Bilen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.