Visual question answering models often encounter challenges related to data biases and exhibit limited performance in specialized domains such as cultural heritage, where recognizing fine-grained textures is crucial. In this study, we propose a novel spatial-frequency invariant semantic learning model designed to overcome these limitations. By incorporating frequency-domain features as an additional modality and employing invariant feature learning techniques, our model effectively reduces bias without relying on external datasets. The proposed model extracts invariant representations across textual, spatial, and frequency domains, thereby filtering out spurious correlations. Comprehensive experiments on benchmark datasets, including VQA-CP v2 and GQA-OOD, demonstrate that our model achieves state-of-the-art results. Furthermore, our model exhibits enhanced robustness when applied to cultural heritage datasets, proficiently handling complex visual textures and multimodal reasoning tasks. This model enhances the capabilities of visual question answering systems in identifying artistic materials and techniques, providing a robust solution tailored to domain-specific applications.
Visual Question Answering (VQA) systems have achieved impressive performance with the rise of large-scale vision–language models (VLMs). However, these models remain vulnerable to multiple forms of multimodal bias, severely limiting their robustness and generalization. Existing debiasing techniques mainly depend on post hoc evaluation or architectural modifications, while recent prompt-learning-based methods reveal new opportunities for aligning downstream tasks with pretrained models. In this work, we propose a unified prompt-driven debiasing framework that integrates generative prompt learning and a fuzzing-based bias correction mechanism. The generative prompt component reformulates VQA as a cloze-style masked prediction problem, leveraging pretrained language priors to improve semantic grounding. Meanwhile, the fuzzing-based module actively constructs unexpected test samples during training and employs a reflection mechanism to correct biased predictions in-loop, yielding inference-time robustness without additional test-time components. Extensive experiments on VQA-v2, VQA-CP, VQA-CE, GQA-OOD, and VQA-VS demonstrate that the proposed framework significantly improves both in-distribution (ID) accuracy and out-of-distribution (OOD) robustness, outperforming existing prompt-only or data-augmentation-only debiasing methods.
Yali Fan, Gangyu Huang, Qiwen Lu et al.· Multimodal Technologies and...· 0 citations