Skip to content

Author

Jianfei Chen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

This work introduces TD-V2A, which leverages temporal differences (TD) as the key representation that distinguishes V2A from I2A, enriching visual conditioning with minimal architectural modification and significantly improves end-to-end V2A generation quality, even outperforming dedicated V2A representations such as contrastive audio-visual pretraining.

Zehua Chen, Junyou Wang, Yuxuan Jiang et al. · 0 citations