How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
SciFigBench is introduced, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty and proposes the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist mi...