Preprint
Jul 2026
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception
BVS bridges prior guidance with posterior correction: it utilizes an early-stop attention rollout of MLLM to construct reasoning-aware priors, while employing a scale-aware non-stationary kernel and GP-UCB to dynamically rectify noise and recover missing information in the prior through iterative local observations.
Geng Li, Yuxin Peng
· 0 citations