EgoTSR++: Egocentric Spatiotemporal Reasoning for Task Progress Understanding
Vision-Language Models (VLMs) have advanced rapidly in static visual understanding, yet remain unreliable when judging how an egocentric task is progressing. Given a task instruction and two visual observations, a model should determine which state is closer to the goal by analyzing task-relevant object configurations...