Preprint
Jul 2026
Token-Based Affordance Grounding with Large Vision-Language Models
TokAG, a zero-shot affordance grounding framework that exploits the token-level semantic-spatial signals in LVLMs to localize action-relevant regions without external supervision, and introduces a spatial-aware token-selection mechanism to systematically evaluate each output token.
Seung Il Lee, Qinqian Lei, Daguang Xu et al.
· 0 citations