Skip to content

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

Jul 2026 · arXiv.org · Vol abs/2607.18827 · 0 citations · 60 references
Computer Science

TL;DR

This work proposes a framework that leverages text-driven object discovery to localize potential gaze candidates, with a gaze-guided selection module to pinpoint the intended target from the candidate objects, and introduces Gradient-Informed Selection Tuning (GIST) to selectively update parameters most relevant to a given class vocabulary.

Abstract

Gaze Object Prediction (GOP) aims to localize and recognize the objects humans attend to, a task crucial for understanding human-centric interactions. However, existing methods are typically trained under a closed-vocabulary paradigm with a fixed label space and evaluated on scene-specific datasets, limiting their applicability to real-world scenarios where gaze targets often follow a long-tail distribution or belong to unseen categories. To address this gap, we introduce Diverse Scenes for Gaze object prediction (DiSG), a benchmark containing 86 in-the-wild categories that facilitates the evaluation of Open-Vocabulary GOP (OVGOP). Building on DiSG, we propose a framework that leverages text-driven object discovery to localize potential gaze candidates, with a gaze-guided selection module to pinpoint the intended target from the candidate objects. Furthermore, to better capture semantic knowledge across diverse in-the-wild categories, we introduce Gradient-Informed Selection Tuning (GIST) to selectively update parameters most relevant to a given class vocabulary. Extensive experiments demonstrate that our proposed model performs effectively in open-vocabulary settings and also outperforms existing methods in the conventional closed-vocabulary setting. The benchmark and code is available at https://github.com/sensniu/ovgop.

View source

Similar papers

Conference Aug 2026

Open-Vocabulary Visual Relationship Detection Via Vision-Language Models And Attention Mechanisms

Traditional Visual Relationship Detection (VRD) systems are strictly bound to predefined, closed-set vocabularies, limiting their practical application in real-world environments where object interactions are highly diverse. Moving towards an open-vocabulary setting (OV-VRD) provides more flexibility but introduces the...

N. Nguyen, Trung Tran · 0 citations
Preprint Sep 2026

From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation

This work introduces the first training-free Gaze Target Agent (GTA) for gaze-guided reasoning across tasks such as gaze target prediction, attention localization, and object identification by leveraging pretrained vision-language models, augmenting them with visually guided prompts, and employing a memory-based retrie...

Shayan Nasiriboukani, Sara Atito, Mohammad Nezamipour et al. · 0 citations
Preprint Sep 2026

OpenVAM: Open-World Visual Attention Modeling with VLMs

Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elem...

Kiana Hooshanfar, A. Kazerouni, Alireza Hosseini et al. · 0 citations
Book Oct 2026

Where Might You Look? Task-Driven Gaze Prediction for Attack-Surface Estimation in Mixed Reality

We propose Gaze-Task Search (GTS), a model which predicts task-driven human gaze as an estimate of the visual attack surface of Mixed Reality (MR) scenes. Because MR by design alters a user's view of the world, a natural security question follows: Which regions of the scene are most likely for an adversary (with access...

Michael Sprintson, Jonathan R. Williford, Jennifer Xu · 0 citations
Preprint Sep 2026

Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide fram...

Yi-Han Zhou, Rui Yan, Ming-Cong Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.