Skip to content
Preprint

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.

Abstract

Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce SafeSceneReason, a multimodal industrial-safety reasoning benchmark and companion training corpus that connects workplace scenes with knowledge from occupational accident investigations. SafeSceneReason combines two complementary data-construction pipelines. The scene-centric pipeline converts annotated workplace images into executable safety scene graphs and generates deterministic answers through program execution over objects, relations, and safety rules. The report-centric pipeline extracts figures and contextual evidence from accident reports and constructs multimodal questions using evidence graphs, explicit information boundaries, multi-step reasoning paths, and iterative verification. The resulting resource contains 110,581 verified scene-centric question--answer pairs and 13,114 refined report-centric question--answer pairs, covering perception, spatial and quantitative reasoning, compliance assessment, evidence synthesis, causal analysis, and mitigation-oriented decision making. Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.

View source

Similar papers

Conference Open access 2026

Mitigating Safety Context Amnesia in Multimodal Reasoning Models via Intent-Guided Safety Reasoning

Intent-Guided Safety Reasoning (IGSR), an inference-time defense that operates without modifying target model parameters, is proposed, which improves defense success rates by over 62% compared to baselines, while largely preserving task utility.

Xiyao Dong, Guangsheng Cheng, Yilong Chen et al. · 0 citations
Preprint Jul 2026

SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

Using SAFERELBENCH to evaluate seven open- and closed-source VLM-driven embodied agents, this work finds a substantial gap between task success and process-level safety compliance, and shows that safe embodied intelligence requires not only stronger perception and planning, but also reliable reasoning about how object relations shape risk during interaction.

Huaigang Yang, Ya Li, Min Ren et al. · 1 citation

Retrieval-Augmented Multimodal Large Language Models for Visual Question Answering of Construction Occupational Health and Safety Hazards

A visual knowledge enhancement framework for construction OHS visual question answering (VQA) based on multimodal large language models (MLLMs) and retrieval-augmented generation (RAG) is proposed, augmenting managerial capacity for reliable and objective OHS hazard prevention.

Yang Liu, Luping Li, Xing Su et al. · 0 citations
Conference Aug 2026

SafetyBuddy: A Multimodal LLM-Based Safety Intelligence Platform for Regulatory Compliance in Process Industries

The problems addressed in this paper are platform architecture, multimodal data fusion and LLM grounding mechanism in the Nigerian industrial system in addition to evaluation metrics, security control, human-in-the-loop validation of safety alerts and the limitations to real-world deployment.

E. C. Ashinze · 0 citations
Preprint Jul 2026

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

This work evaluates several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether an explicit distinction between hazard and anomaly changes model behavior, and shows that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure.

M. Indukuri, Mohammad Eskandari, Sree Nitya Kollu et al. · 0 citations
Preprint Aug 2026

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

CAViAR exposes a practical Perception--Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule-relevant responsibility categories in safety-critical driving scenarios, and all models degrade sharply on accident type and responsibility reasoning.

Sparsh Garg, Yi-Wen Chen, Vijay Kumar et al. · 0 citations