Skip to content
Conference

Probabilistic Scene Graph Prompting: Uncertainty-Aware Structured Reasoning in Multimodal LLMs

Mar 2026 · 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW) · pp. 1695-1702 · 1 citation · 24 references

Abstract

Scene graphs provide structured visual knowledge that can enhance multimodal large language models (MLLMs) for tasks like visual question answering and image captioning. However, existing approaches inject scene graphs as deterministic, hard prompts, ignoring the inherent uncertainty in visual perception, leading to overconfident and sometimes hallucinated outputs. We propose Probabilistic Scene Graph Prompting (PSGP), a framework that models scene graph generation as a distribution over plausible graphs and encodes this uncertainty into soft, continuous prompt tokens that condition the MLLM. By propagating perceptual uncertainty from detection to language generation, PSGP produces more accurate, faithful, and better-calibrated responses, especially in ambiguous visual scenarios. Experiments on GQA, Visual Spatial Reasoning, and a new Ambiguous-GQA benchmark show that PSGP outperforms strong baselines: including LLaVA-1.5, BLIP-2, and SG-LLaVA, in accuracy, faithfulness, and calibration, while maintaining computational efficiency. Our work establishes a principled pathway toward uncertainty-aware, structured multimodal intelligence.

View source