This paper presents a review and practical study of AI-based phishing detection. It explains how machine learning can be used to identify phishing websites using URL and webpage features. The paper discusses common machine learning models such as Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, and Neural Networks. It also explains the main steps of building a phishing detection system, including data collection, preprocessing, feature extraction, model training, and evaluation. The paper discusses the challenges and limitations of machine-learning-based phishing detection and suggests areas for future work.
Navneet Yadav· Zenodo (CERN European Organi...· 0 citations
This paper proposes a novel neuro-symbolic reasoning framework based on Hierarchical Bayesian Networks (HBNs). The core challenge in neuro-symbolic reasoning lies in effectively integrating the pattern recognition capabilities of neural networks with the structured reasoning capabilities of symbolic systems. Traditional approaches often struggle with knowledge representation and the ability to handle uncertainty. This work addresses these limitations by constructing a system where HBNs are used to represent and reason about knowledge hierarchically. Neural networks are employed to learn specific features and relationships within the HBN structure, while the Bayesian network provides a framework for probabilistic inference and reasoning under uncertainty. The hierarchical structure enables the system to decompose complex problems into smaller, more manageable sub-problems, improving both accuracy and interpretability. This approach demonstrates the potential to create more robust and explainable AI systems capable of handling complex reasoning tasks. The key contribution is the specific application of HBNs for this integration, moving beyond simple neural-symbolic hybrids. We outline the architecture, learning process, and inference mechanisms, focusing on the benefits of a structured knowledge representation.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
This paper investigates the application of Explainable Artificial Intelligence (XAI) to enhance system security. Traditional security systems often rely on opaque "black box" AI models, hindering effective threat detection and defense. This work proposes a system architecture leveraging XAI techniques to provide interpretable insights into AI-driven security decisions. The core claim is that building an XAI-based system for security threat detection and defense will improve both the efficiency and reliability of security measures. The proposed mechanism utilizes XAI to elucidate the reasoning behind AI's judgments, thereby facilitating the identification of vulnerabilities and potential threats. Specifically, we explore methods for generating explanations that highlight critical factors influencing security decisions, allowing human analysts to validate, refine, or override AI recommendations. This approach addresses the limitations of current black-box AI security systems by incorporating human understanding and control, ultimately leading to a more robust and trustworthy security posture. The research contributes to a paradigm shift in system security, moving beyond solely relying on AI's predictive capabilities to actively incorporating human expertise within the decision-making process.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
Current Explainable AI (XAI) techniques frequently generate explanations that are merely post-hoc justifications for model predictions, lacking a deep understanding of the model's reasoning process. This paper proposes a novel approach to XAI based on the construction of Causal Bayesian Networks (CBNs). CBNs are employed to explicitly model the causal relationships between input features and the model's output, offering a more robust and interpretable explanation. Unlike existing XAI methods which often rely on correlations, our approach leverages causality to provide a truly grounded understanding of how the model arrives at its decisions. The method is presented with a detailed theoretical framework and outlines the steps involved in constructing and utilizing CBNs for explainability. This research addresses a critical limitation in current XAI by moving beyond correlation-based explanations towards a causal understanding, ultimately leading to more reliable and trustworthy AI systems.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This is the initial release of "dogo-tutor," an AI tutor designed for coding education. This release is created to archive the repository and obtain a DOI via Zenodo. Key Features Provides step-by-step hints (problem clarification, strategy, logic structuring, and partial code) rather than direct answers. Designed to offer deeper hints only when students explain "what they tried and what happened" in their own words. Automatically generates review notes and short quizzes (/review) based on the last 24 hours of dialogue history, focusing on concepts the student struggled with. Features a self-reflection tool (/log) that allows students to objectively track their dependency on deep hints. Employs a safe guidance design that strictly references URLs from actual lecture materials and official documentation specified by the instructor. Recent Changes Revamped the README for users and added an English summary. Adjusted the icon design for the VS Code extension. Updated various configuration files (e.g., package.json) for public release.
KAZUHIRO JINMA· Zenodo (CERN European Organi...· 0 citations
Coexisting safely with unpredictable Artificial Intelligence (AI) remains a foundational challenge for contemporary AI ethics. This paper proposes “stable otherness” as a relational framework that re-centers alignment on sociotechnical structures. Rather than treating AI as an autonomous conscious agent, we define AI as an inanimate “pseudo-otherness”—a mechanical externalization of human reflexive cognition. Drawing on a heuristic natural-historical scaffolding of interspecies relations, we analyze how relational predictability, interpretability, and response stability emerge. We show that AI’s unpredictability, unlike biological “wildness,” stems from structural limits including symbol-grounding deficits and next-token prediction dynamics. Operationalizing stable otherness through the triad of Explainability, Alignment Stability, and Safe-by-Design, we discuss implications for governance and argue that claims of AI rights may constitute a category error, thereby re-anchoring ethical responsibility in human design. This framework directly contributes to AI ethics, information ethics, and the philosophy of technology, providing a clear evaluative structure for alignment and governance.
Fumio Miyata· Zenodo (CERN European Organi...· 0 citations
The rapid convergence of artificial intelligence (AI) workloads, high-density computing, advanced liquid and air cooling, increasingly distributed electrical architectures, and stringent availability requirements is transforming the data centre from a predominantly passive infrastructure environment into a highly dynamic cyber-physical system. Conventional data centre infrastructure management (DCIM) platforms remain largely dependent on fragmented telemetry, threshold-based alarms, static visualization, and human interpretation, limiting their ability to continuously correlate physical infrastructure behaviour with computational workload, commissioning state, operational risk, and lifecycle constraints. This paper proposes AIDCDT-X, an AI-native autonomous data centre digital twin framework that extends the conventional digital twin concept from visualization and monitoring toward predictive, context-aware, and governance-controlled infrastructure intelligence. The proposed architecture evolves the original five-layer AIDCDT model into an integrated cyber-physical intelligence stack comprising physical asset representation, synchronized multi-domain telemetry, contextual digital-twin state estimation, predictive and prescriptive intelligence, decision orchestration, and governance/audit control. The framework establishes a continuous information loop connecting building management systems (BMS), electrical power management systems (EPMS), DCIM, information-technology telemetry, environmental sensing, commissioning records, asset configuration, and operational events. A central contribution is a lifecycle-aware twin maturation model linking Design, Commissioning, Operational Readiness, and Steady-State Operation. Rather than treating commissioning documentation as a terminal project artifact, the proposed approach converts Level 1–5 commissioning tests, integrated systems testing, defect registers, burn-in observations, and as-built configuration data into structured training and calibration evidence for the operational twin. This creates a facility-specific intelligence baseline capable of continuously updating its representation of infrastructure health, thermal behaviour, electrical resilience, capacity margin, and operational risk. The paper further introduces a proposed multi-objective cyber-physical optimization formulation in which latency, energy efficiency, thermal stability, availability, predictive uncertainty, maintenance risk, lifecycle cost, and governance constraints are jointly evaluated. A bounded-autonomy principle is incorporated to explicitly separate prediction from authorized physical action, thereby preventing unrestricted AI control of mission-critical infrastructure. The resulting framework provides a pathway toward explainable and auditable autonomous operations while retaining human authorization at safety-critical decision boundaries. Unlike a conventional conceptual digital twin, AIDCDT-X is designed to support measurable validation through predictive lead time, anomaly-detection performance, commissioning defect detection, decision latency, false-positive and false-negative rates, availability impact, energy efficiency, and human-override metrics. The present work establishes the architecture and analytical methodology; it does not claim empirical performance from a live facility. Future experimental validation will therefore focus on a controlled pilot or high-fidelity digital-twin environment using facility-specific commissioning and operational datasets.
M. Rizwan Yasin· Zenodo (CERN European Organi...· 0 citations
Chemical safety in the Republic of Korea is governed by five acts falling under the purview of the Ministry of Climate, Energy, and Environment (MCEE): the Act on Registration and Evaluation of Chemicals (K-REACH Act), the Chemicals Control Act (CCA), the Environmental Health Act, the Environmental Damage Relief Act, and the Consumer Chemical Products and Biocides Safety Act.These five acts manage substantial datasets but operate on largely independent data systems, classification schemes, and identifiers, which limits cross-act risk prediction and timely policy responses.Advances in artificial intelligence (AI) and knowledge graph technologies suggest a possible paradigm shift, but realizing this potential calls for a redesign of data infrastructure and institutional frameworks tailored to the five-act structure.Drawing on international best practices, including the European Union (EU) One Substance One Assessment (OSOA) package, the EU Registration, Evaluation, Authorisation and Restriction of Chemicals (REACH) regulation, the United States Toxic Substances Control Act (TSCA), Data Catalog Vocabulary-Application Profile (DCAT-AP) metadata standards, and ontology-and knowledge graph-based chemical safety studies, as well as on the structural lessons of past data-integration failures such as the 9/11 information silos, a seven-action strategy is proposed: (1) Designating high-value chemical-safety datasets across the five acts; (2) Redesigning data collection around policy questions; (3) Adopting DCAT-based metadata standards; (4) Developing a chemical-safety domain ontology with cross-act bridges anchored by a common chemical identifier; (5) Transforming incident reports into knowledge graphs; (6) Building AI-based early warning and prioritization systems; and(7) Institutionalizing explainable AI.The seven actions form a logical pipeline from data through metadata, ontology, knowledge graph, AI, and explainability, and are intended to be pursued in a stepwise manner.Realizing this transformation will likely require coordination across the five acts within MCEE, legal safeguards, and sustained investment in data curation and explainable AI.
Ho-Hyun Kim, Lim Ho-Ju, Hunjoo Lee· Korean Journal of Environmen...· 0 citations
The construction of health indicators (HIs) for failure prognostics is crucial for monitoring and predicting the health state of multi-component systems. Recently, data-driven methods for HI construction and prognostics have gained increasing attention, both relying heavily on high-quality run-to-failure (RTF) data for accurate analysis and reliable prediction. To foster trust in these systems, integrating explainable AI (XAI) into prognostics and health management (PHM), forming XPHM, is highly recommended. However, current literature reveals several research gaps: a lack of RTF data for multicomponent systems, limited consideration of complex component interactions, and insufficient attention to environmental impacts in constructing explainable HIs for prognostics. To address these challenges, this paper introduces new RTF data for the Tennessee Eastman Process (TEP), capturing multicomponent interactions under various environmental conditions through a systematic simulation methodology. Furthermore, the study analyzes this data, elucidating the direct and indirect relationships between sensor measurements and TEP component states across different working conditions. These findings provide valuable insights for researchers and practitioners in developing both component- and system-level HIs and prognostics. Additionally, we illustrate how to use the provided data and analyzed results to construct HIs, further supporting the development of reliable XAI solutions, thereby improving the reliability of engineering systems.
Duc An Nguyen, Nguyen, Khanh, T P, Kamal Medjaher· Proceedings of the Instituti...· 0 citations
This survey explores the use of Machine Learning (ML) in the field of Computer Algebra (CA), both to optimise existing CA algorithms, and to perform symbolic computation directly. Traditional symbolic methods, while mathematically rigorous, often are computationally expensive thus limiting their use in real-world applications. Recent advances have shown that data-driven techniques can address these limitations by guiding heuristic decisions, selecting optimal algorithms, predicting structural properties of algebraic objects, or even making direct symbolic computations. We take a systematic literature review approach and uncover work in CA applications including cylindrical algebraic decomposition, Gröbner basis computation, symbolic integration, and many more. The survey compares the different ML approaches that have been employed for these tasks, ranging from decision trees to transformers. Issues uncovered by the survey include the lack of benchmark datasets for CA, which hinders the comparison of methods and the generalizability of ML models. The survey identifies the potential for explainable AI tools to help develop trust in decisions, and to drive forward CA research itself.
Uzma Shafiq, Matthew England, Nayyar Zaidi· Pure (Coventry University)· 0 citations
Abstract Effective Pavement Management System (PMS) planning depends on the ability to anticipate both the severity and physical extent of multiple distress types under real-world data constraints. This work introduces an interpretable dual-stage Deep Learning (DL) framework that sequentially classifies distress severity and then predicts the corresponding crack lengths, alligator cracking areas, and pothole quantities. The first stage uses an Artificial Neural Network (ANN) optimized with Focal loss to overcome severe class imbalance, while the second employs an ANN with Huber loss and targeted non-zero weighting to counteract outlier dominance and the zero-inflation inherent in sparse distress records. Data scarcity is addressed through the Synthetic Minority Over-sampling Technique (SMOTE) and Gaussian noise augmentation. To ensure the learned relationships are not merely correlational, an Explainable Artificial Intelligence (XAI) framework combining Shapley Additive Explanations (SHAP), Partial Dependence Plots (PDP), and Individual Conditional Expectation (ICE) curves is integrated, verifying that the model's internal logic adheres to established Mechanistic-Empirical (M-E) pavement science—most notably by recovering high-severity alligator cracking as the dominant antecedent to pothole formation. Beyond prediction quality, the design is substantiated through systematic ablation experiments: a shared-backbone Multi-Task Learning (MTL) variant is shown to degrade regression accuracy due to gradient interference between conflicting loss objectives, while network capacity is empirically scaled, using a bottleneck regularization strategy for data-scarce targets such as pothole area and count, to prevent overfitting without compromising expressiveness. The framework achieves an F1-score of 0.982 for pothole severity, and R² values of 0.954 for alligator cracking area and 0.941 for linear crack length. Paired t-tests confirm the absence of systematic bias, establishing the architecture as a high-fidelity and trustworthy input for pavement maintenance planning within the Long-Term Pavement Performance (LTPP) context.
Mohammad Sedighian-Fard, Amir Golroo, Mahdi Javanmardi et al.· Scientific Reports· 0 citations
Coexisting safely with unpredictable Artificial Intelligence (AI) remains a foundational challenge for contemporary AI ethics. This paper proposes “stable otherness” as a relational framework that re-centers alignment on sociotechnical structures. Rather than treating AI as an autonomous conscious agent, we define AI as an inanimate “pseudo-otherness”—a mechanical externalization of human reflexive cognition. Drawing on a heuristic natural-historical scaffolding of interspecies relations, we analyze how relational predictability, interpretability, and response stability emerge. We show that AI’s unpredictability, unlike biological “wildness,” stems from structural limits including symbol-grounding deficits and next-token prediction dynamics. Operationalizing stable otherness through the triad of Explainability, Alignment Stability, and Safe-by-Design, we discuss implications for governance and argue that claims of AI rights may constitute a category error, thereby re-anchoring ethical responsibility in human design. This framework directly contributes to AI ethics, information ethics, and the philosophy of technology, providing a clear evaluative structure for alignment and governance.
Fumio Miyata· Zenodo (CERN European Organi...· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.