DGEval is introduced, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24, and shows that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context.
Abstract
The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on the NCB Hazcheck e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.
WuYu-EnvLE-Bench is introduced, a benchmark built from real enforcement cases, regulatory standards, and expert review that highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.
Ziliang Yang, Yi Zhang, K. Lin et al.· 0 citations
The increasing complexity of global e-commerce supply chains underscores the need for automated, context-aware risk monitoring systems capable of interpreting large volumes of unstructured information. Although Large Language Models (LLMs) have shown strong performance across natural language processing tasks, their application to real-world supply chain risk detection remains limited. This study presents a novel, manually annotated dataset of 121 business news articles related to five major steel companies, using the Cambridge Risk Taxonomy. Leveraging this dataset, we evaluate two state-of-the-art LLMs in a multi-label risk classification task using few-shot prompting. The results demonstrate that LLMs can approximate human annotation, though challenges persist in detecting domain-specific risks such as Geopolitical threats and in avoiding label overgeneration. Beyond classification, we further assess the capacity of LLMs to generate managerial risk summaries. We show that summaries derived from model-predicted risks exhibit strong semantic alignment to summaries generated from human annotations, highlighting the potential of LLMs to support executive-level risk interpretation. Overall, this study contributes the first publicly available dataset of fine-grained, hierarchical risk annotations in an e-commerce supply chain context and provides empirical evidence on the opportunities and limitations of LLMs for both analytical and narrative forms of automated risk assessment.
Laleh Davoodi, Filip Ginter, Sima Salimi et al.· SN Computer Science· 0 citations
The problems addressed in this paper are platform architecture, multimodal data fusion and LLM grounding mechanism in the Nigerian industrial system in addition to evaluation metrics, security control, human-in-the-loop validation of safety alerts and the limitations to real-world deployment.
E. C. Ashinze· SPE Nigeria Annual Internati...· 0 citations
Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee operational reliability: a misread altitude, a dropped execution condition, or a confused call- sign may score well under standard F1 yet carry sharply asymmetric operational consequences. We study this problem in air traffic control (ATC), where controller-pilot communication demands near-zero error tolerance, and use consequence-aware evaluation to test whether semantic scores misstate operational reliabil- ity. The framework is instantiated in a con- trolled diagnostic ATC benchmark grounded in aviation standards and feedback from 40 air traffic controllers across three countries. Evaluating 8 models, we uncover a system- atic semantic-safety gap: conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics. Risk-aware fine-tuning narrows but does not close this gap, showing that consequence- aware evaluation is a necessary complement to standard NLP metrics before any real safety- critical deployment claim
Yujing Chang, Thinh Pham, Van-Phat Thai et al.· 0 citations
AtmosCoder-Bench is introduced, an execution-grounded benchmark that makes the calculation process visible, and finds that multiple-choice formats inflate measured accuracy by at least 12 percentage points.
Maohao Ran, Chendong Ma, Yanting Zhang et al.· 0 citations
Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64\% accuracy on the Foundation Module, but average accuracy still fell from 84.14\% on easy questions to 42.50\% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.
Yi Zhang, Hongyang Wang, Zheng Hao Leong et al.· 0 citations