Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.
Abstract
Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce SafeSceneReason, a multimodal industrial-safety reasoning benchmark and companion training corpus that connects workplace scenes with knowledge from occupational accident investigations. SafeSceneReason combines two complementary data-construction pipelines. The scene-centric pipeline converts annotated workplace images into executable safety scene graphs and generates deterministic answers through program execution over objects, relations, and safety rules. The report-centric pipeline extracts figures and contextual evidence from accident reports and constructs multimodal questions using evidence graphs, explicit information boundaries, multi-step reasoning paths, and iterative verification. The resulting resource contains 110,581 verified scene-centric question--answer pairs and 13,114 refined report-centric question--answer pairs, covering perception, spatial and quantitative reasoning, compliance assessment, evidence synthesis, causal analysis, and mitigation-oriented decision making. Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.
Intent-Guided Safety Reasoning (IGSR), an inference-time defense that operates without modifying target model parameters, is proposed, which improves defense success rates by over 62% compared to baselines, while largely preserving task utility.
Xiyao Dong, Guangsheng Cheng, Yilong Chen et al.· Annual Meeting of the Associ...· 0 citations
Using SAFERELBENCH to evaluate seven open- and closed-source VLM-driven embodied agents, this work finds a substantial gap between task success and process-level safety compliance, and shows that safe embodied intelligence requires not only stronger perception and planning, but also reliable reasoning about how object relations shape risk during interaction.
A visual knowledge enhancement framework for construction OHS visual question answering (VQA) based on multimodal large language models (MLLMs) and retrieval-augmented generation (RAG) is proposed, augmenting managerial capacity for reliable and objective OHS hazard prevention.
Yang Liu, Luping Li, Xing Su et al.· Journal of Management in Eng...· 0 citations
The problems addressed in this paper are platform architecture, multimodal data fusion and LLM grounding mechanism in the Nigerian industrial system in addition to evaluation metrics, security control, human-in-the-loop validation of safety alerts and the limitations to real-world deployment.
E. C. Ashinze· SPE Nigeria Annual Internati...· 0 citations
This work evaluates several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether an explicit distinction between hazard and anomaly changes model behavior, and shows that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure.
M. Indukuri, Mohammad Eskandari, Sree Nitya Kollu et al.· 0 citations
CAViAR exposes a practical Perception--Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule-relevant responsibility categories in safety-critical driving scenarios, and all models degrade sharply on accident type and responsibility reasoning.
Sparsh Garg, Yi-Wen Chen, Vijay Kumar et al.· 0 citations