Open-Vocabulary Visual Relationship Detection Via Vision-Language Models And Attention Mechanisms
Abstract
Traditional Visual Relationship Detection (VRD) systems are strictly bound to predefined, closed-set vocabularies, limiting their practical application in real-world environments where object interactions are highly diverse. Moving towards an open-vocabulary setting (OV-VRD) provides more flexibility but introduces the challenge of accurately matching unseen relationships with unconstrained text queries. In this paper, we present an OV-VRD framework that directly embeds pre-trained Vision-Language Models (VLMs), specifically CLIP, into a translation embedding architecture. By replacing conventional static word embeddings with CLIP’s joint semantic space, the model can effectively generalize to novel predicates. To further capture the nuanced contextual interactions between subjects and objects, we design a customized multi-head attention module. Crucially, this attention block is modified to exclude the standard linear projection layer, thereby preserving the integrity of the original VLM features. Experimental results on the Stanford VRD and VG50 datasets confirm the effectiveness of this design. Notably, the integration of the attention mechanism alone yields a 3.35% absolute improvement in Recall@50 on the VRD benchmark compared to the baseline CLIP configuration establishing a strong and highly competitive baseline for open-vocabulary relationship reasoning.