Skip to content
Book

Where Might You Look? Task-Driven Gaze Prediction for Attack-Surface Estimation in Mixed Reality

Oct 2026 · Proceedings of the Second Workshop on Enhancing Security, Privacy, and Trust in Extended Reality · 0 citations · 10 references

Abstract

We propose Gaze-Task Search (GTS), a model which predicts task-driven human gaze as an estimate of the visual attack surface of Mixed Reality (MR) scenes. Because MR by design alters a user's view of the world, a natural security question follows: Which regions of the scene are most likely for an adversary (with access to the rendering layer) to attack in order to distract or prevent the user from accomplishing a task? Answering this requires predicting where a human tends to look while searching, which lies between the well established domains of predicting natural saliency and predicting fixations for task-based search. Existing saliency models ignore the task, while goal-directed search models depend on category-limited eye-tracking data to predict fixations. GTS instead extracts cross-attention from a frozen vision-language diffusion model into a bottom-up saliency backbone, introducing world model knowledge to the saliency task, reducing the reliance on task-specific eye-tracking data for training. Rather than predicting exact fixations, we produce a heatmap of saliency across the scene. On the VIU dataset, GTS outperforms a leading saliency model, UNISAL, on five of six metrics for both target-present and target-absent images, improving normalized scanpath saliency by 6% in the hardest, target-absent setting. We frame its predictions as an attack-surface estimate for MR threat modeling.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.