Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning
Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the resulting activation shifts may accumulate across decoding steps, progr...