This work proposes a framework that leverages text-driven object discovery to localize potential gaze candidates, with a gaze-guided selection module to pinpoint the intended target from the candidate objects, and introduces Gradient-Informed Selection Tuning (GIST) to selectively update parameters most relevant to a given class vocabulary.
Abstract
Gaze Object Prediction (GOP) aims to localize and recognize the objects humans attend to, a task crucial for understanding human-centric interactions. However, existing methods are typically trained under a closed-vocabulary paradigm with a fixed label space and evaluated on scene-specific datasets, limiting their applicability to real-world scenarios where gaze targets often follow a long-tail distribution or belong to unseen categories. To address this gap, we introduce Diverse Scenes for Gaze object prediction (DiSG), a benchmark containing 86 in-the-wild categories that facilitates the evaluation of Open-Vocabulary GOP (OVGOP). Building on DiSG, we propose a framework that leverages text-driven object discovery to localize potential gaze candidates, with a gaze-guided selection module to pinpoint the intended target from the candidate objects. Furthermore, to better capture semantic knowledge across diverse in-the-wild categories, we introduce Gradient-Informed Selection Tuning (GIST) to selectively update parameters most relevant to a given class vocabulary. Extensive experiments demonstrate that our proposed model performs effectively in open-vocabulary settings and also outperforms existing methods in the conventional closed-vocabulary setting. The benchmark and code is available at https://github.com/sensniu/ovgop.
Traditional Visual Relationship Detection (VRD) systems are strictly bound to predefined, closed-set vocabularies, limiting their practical application in real-world environments where object interactions are highly diverse. Moving towards an open-vocabulary setting (OV-VRD) provides more flexibility but introduces the...
N. Nguyen, Trung Tran· International Conference on...· 0 citations
This work introduces the first training-free Gaze Target Agent (GTA) for gaze-guided reasoning across tasks such as gaze target prediction, attention localization, and object identification by leveraging pretrained vision-language models, augmenting them with visually guided prompts, and employing a memory-based retrie...
Shayan Nasiriboukani, Sara Atito, Mohammad Nezamipour et al.· 0 citations
Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elem...
Kiana Hooshanfar, A. Kazerouni, Alireza Hosseini et al.· 0 citations
We propose Gaze-Task Search (GTS), a model which predicts task-driven human gaze as an estimate of the visual attack surface of Mixed Reality (MR) scenes. Because MR by design alters a user's view of the world, a natural security question follows: Which regions of the scene are most likely for an adversary (with access...
Michael Sprintson, Jonathan R. Williford, Jennifer Xu· Proceedings of the Second Wo...· 0 citations
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide fram...
Yi-Han Zhou, Rui Yan, Ming-Cong Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.