JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting
Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video generation models produce visually plausible predictions but require expensive end-to-end training and often entangle scene understanding with image synthesis. In this wor...