Preprint
Jul 2026
ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
Spatio-Temporal Token Veto is proposed, which leverages the ability to observe all token positions at each diffusion step and vetoes temporally unstable tokens via second-order Taylor prediction of confidence dynamics and filters weakly grounded tokens using image-attention mass, swapping them with safer candidates.
Keuntae Kim, Beomseok Lee, Hyunwoo Kim et al.
· 0 citations