This study proposes a lightweight Natural Language Processing (NLP) pipeline to automatically classify the primary cause of major accidents using the European eMARS database and shows that Word2Vec+SVM provides the strongest and most stable baseline on the full labelled set, while SBERT performance improves markedly under higher label fidelity.
This study presents a novel, AI-assisted approach to industrial reliability analysis that integrates Failure Modes, Effects and Criticality Analysis (FMECA) and the Failure Reporting, Analysis, and Corrective Action System (FRACAS) with Natural Language Processing (NLP). We developed an algorithm that automates and streamlines the analysis of equipment field-failure reports and other unstructured maintenance records. The proposed framework combines unsupervised clustering to identify recurring equipment failure modes with a supervised Support Vector Machine (SVM) classifier with a Radial Basis Function (RBF) kernel to categorize equipment field reliability reports by failure mode and underlying mechanism at scale. Using an train–test split, the proposed model achieved accuracy on the test dataset, indicating effective generalization to unseen maintenance reports. The Confusion Matrix metrics across all classes showed true positive rate (TPR) (or sensitivity) of 0.91, indicating the model’s strong ability to correctly identify positive samples. The false positive rate (FPR) averaged 0.03 across all classes, demonstrating excellent specificity (true negative rate, TNR of 0.97). Operationally, the methodology reduces the resource-intensive manual work required to prepare, interpret, and process FRACAS reports, thus enabling timelier, data-driven equipment reliability analysis. Overall, the study demonstrates the feasibility and benefits of using AI-assisted reliability tools that balance automation with human expertise through a human-in-the-loop approach.
Esther Yu, Guangjiang Cao, Y. Khalil et al.· Journal of Research in Engin...· 0 citations
Learning from incidents is a cornerstone of occupational safety risk management, especially in high-risk industrial sectors. Incident databases support this process by collecting records that describe the causes, dynamics, and consequences of adverse events. However, these databases largely rely on unstructured textual narratives, which limits systematic analysis and the translation of learned lessons into effective preventive actions. This paper focuses on the iron and steel industry and presents an analysis pipeline supported by Large Language Models (LLMs) for extracting, synthesising, and structuring information from two major incident databases: the U.S. OSHA database and the French ARIA database. Relevant records were selected using industry classification codes and pre-processed to harmonise terminology, normalise information fields, remove duplicates, and manage multilingual content. Within a human-in-the-loop framework, LLMs were used to identify critical occupational risk scenarios, characterise them in terms of frequency and severity, and derive prevention and risk mitigation measures structured according to the ISO 45001 hierarchy of controls. Eight critical scenarios were identified and subsequently validated and refined by safety experts from the steel industry. Quantitative analysis identified point-of-operation machinery and load handling as the most frequent scenarios, while confined spaces and high-energy events exhibited disproportionate severity and lethality. The results demonstrated how LLM-supported approaches can enhance learning from incidents by transforming large volumes of heterogeneous narrative data into a traceable, expert-validated knowledge base that supports hazard identification, risk assessment and management, and continuous improvement in high-risk industrial environments.
Giuseppe Tomasoni, F. Marciano, P. Cocca et al.· AHFE International· 0 citations
Knowledge Management Systems (KMS) are required to organize and assign meaning to huge amounts of organizational knowledge that are largely in the form of unstructured text. Natural Language Processing (NLP), and more immediately methods of text categorization, has been one of the principal enabler technologies to enable KMS to be simpler by helping to automatically categorize documents, enhance searching for information, and assist in decision-making. This paper offers an outline of the evolution of NLP-based text classification methods from initial machine learning methods such as Naïve Bayes and Support Vector Machines to current sophisticated deep learning algorithms such as Convolutional Neural Networks, Recurrent Neural Networks, and Transformers. We offer real-world industry use cases, issues of scalability, explainability, and ethics and encapsulate research areas of existing gaps. The findings underscore the enormous potential of NLP text classification to assist the effectiveness and efficiency of knowledge management (KM) activities.
Unknown authors· Engineering and Technology J...· 0 citations
Open-Source Intelligence (OSINT) can be considered a crucial part of the present-day Cyber Threat Intelligence (CTI) due to delivering prompt information about the adversary activity using publicly accessible reporting and analysis. Nonetheless, the conversion of unstructured OSINT stories into structured forms like the MITRE ATT&CK model is a highly manual and subjective task. The proposed paper offers an automated solution to the problem of categorizing the threat descriptions based on OSINT into the MITRE ATT&CK techniques with a sophisticated model based on transformer and natural language processing. The suggested framework has combined OSINT preprocessing, threat behavior extraction, semantic representation learning and multi-label ATT&CK techniques classification with confidence-aware outputs. Large-scale experiments on a wide OSINT corpus show that the proposed method is far more effective compared to the ones that rely on keyword parameters and conventional machine learning baselines, especially when there are missing and imprecise threat specifications. Findings indicate greater accuracy, retrieval, and strength of a vast variety of ATT&CK methods, including categories of low density. This work is a step in the right direction by facilitating scalable and standardized ATT&CK mapping of noisy OSINT data, thus leading to more effective CTI automation and decision support to the security operations of a practitioner.
Ramesh Kumar Sharma, Dharmendra Kumar Singh, Rajesh P. Barnwal· International journal of com...· 0 citations
This study presents an automated knowledge extraction and prediction system using the advancements in Artificial Intelligence (AI) tools, referred to as APEX-LLM, which is a scalable, domain-independent system which can be customized and applied to health, financial and business sectors, and education.
This paper analyzes metadata from Ecuador's SOCE, with particular emphasis on participant comments generated during the pre-contractual phase to propose a hybrid modeling framework that integrates unsupervised clustering and supervised classification within a natural language processing (NLP) pipeline to uncover latent patterns and detect potentially irregular procurement processes.
Bryan Torres, Daniel Riofrío, J. Vega-Sánchez et al.· 0 citations