Category
explainable ai
349 papers
A multimodal cross-modal explainable retrieval-augmented generation framework for hallucination-grounded visual question answering
Recent multimodal large language models (MLLMs) have achieved remarkable performance on visual question answering (VQA) and multimodal reasoning tasks. But one overlooked failure case stubbornly persists: when answers require simultaneous synthesis of evidence from natural images, free-form text, and structured tables, present models suffer from trimodal hallucination—generating content unattributed to any of the input modalities. Prior approaches to hallucination mitigation focus on image-text pairs, and existing retrieval-augmented generation (RAG) systems for multimodal settings do not readily include supporting evidence and often lack modality-attributed explainability. In this paper, we present MXRAG (Multimodal Cross-Modal Explainable Retrieval-Augmented Generation), a new approach to this trimodal evidence problem with three key innovations that work in concert: (1) a Trimodal Evidence Retriever (TER) that retrieves image patches, text passages, and table rows jointly using a shared semantic manifold; (2) a Cross-Modal Attribution Network (CMAN) that computes fine-grained, token-level attribution scores inline during generation, mapping each generated token to supporting evidence from all three modalities; and (3) a Hallucination-Aware Constrained Decoding (HACD) strategy that penalises generation steps with attribution entropy above a calibrated threshold, suppressing unsupported factual tokens at inference time. CMAN is trained with novel cross-modal attribution and modality-coherence losses using token-level gold annotations; HACD requires no additional training and is calibrated per dataset on the validation split. We cast joint retrieval-generation as a constrained variational problem over a trimodal evidence space and introduce MMTabQA, a new trimodal VQA benchmark derived from WikiTableQuestions and MSCOCO with 12,847 instances and token-level attribution labels. Evaluation on four benchmarks (MMTabQA, WebSRC, ChartQA, MIMIC-CXR-VQA) shows that MXRAG achieves state-of-the-art. +21.4% points (p.p.) exact match accuracy, − 38.6% relative reduction in hallucination rate (− 12.1 p.p. absolute) versus the best multimodal RAG baseline, and 89.4% modality coherence score. Ablation experiments confirm the contribution of each component, with CMAN providing the largest accuracy gain (+ 14.2%) and HACD the largest hallucination reduction (− 23.1%). A supplementary human evaluation on 200 instances confirms that the entropy-based hallucination metric tracks human judgement (human-judged HR: 21.3% vs. metric HR: 19.1%). MXRAG advances the state of the art for reliable, interpretable, evidence-based multimodal AI.
Creating Educational Documentation with AI
T his chapter covers the use of all the artificial intelligences (AIs) explained so far (except Brisk Teaching), to demonstrate how teachers can combine them to create any type of educational document, regardless of its application or subject.
The Gut–Brain Connection and the Influence of the Microbiome on Cognition
Cognitive health is not governed by the brain in isolation; it is actively co-regulated by the gut–brain–microbiome axis. This work argues that microbial composition and function are not passive correlates of brain health but causal drivers of neuroplasticity, neuro-inflammation, and long-term cognitive resilience. Trillions of gut microbes generate neuroactive metabolites, regulate immune signalling, and maintain intestinal barrier integrity. When this system is disrupted, most often by low-fibre diets, chronic stress, poor sleep, or indiscriminate medication use, the inflammatory signalling escalates, neurogenesis declines, and vulnerability to neurodegenerative disease increases. Traditional statistical approaches struggle to capture these effects because microbiome–brain interactions are nonlinear, individualized, and temporally dynamic. Artificial intelligence and machine learning change this landscape. By integrating metagenomic, clinical, and lifestyle data, AI models can move beyond surface-level associations to simulate gut–brain interactions and predict individual responses to dietary, probiotic, and behavioural interventions. Recent deep-learning studies using stool metagenomics to predict Parkinson’s disease risk years before clinical onset illustrate both the promise and urgency of this approach. Yet prediction alone is insufficient. The small cohort sizes, population bias, and weak causal inference limit model reliability. Progress depends on hybrid frameworks that couple machine learning with longitudinal sampling, mechanistic validation, and controlled intervention studies. Precision brain health will not replace foundational lifestyle practices, but it can finally explain why they work and for whom, enabling proactive prevention rather than reactive treatment.
Reach audiences
Advertise in front of researchers, engineers, and readers.
V3 Architecture — Visual Atlas of the Periodic Table: 118 Elements as Vortex Clusters (AI-Generated Illustrations)
V3 ARCHITECTURE — COMPLETE PERIODIC TABLE PACKAGECode + 118 Visual Plates + Geometry of Matter + Nuclear Stability Simulation This deposit contains the complete V3 Architecture package for the periodic table,including: 1. ADA SPARK SOURCE CODE (100% GNATprove proof obligations) - V3_Unified.ads/adb : Complete V3 Architecture with geometry, radiation, and 118 elements - V3_Materials_Classifier.adb : Reclassification of 200 materials by phase - V3_Nuclear_Stability.adb : Kinetic simulation of nuclear stability from stable to rupture - V3_Constants, V3_Geometry, V3_Radiation : Full V3 libraries 2. 118 VISUAL PLATES (AI-Generated Illustrations) - Each element represented as a vortex cluster with: - Protonic vortices (orange-red toroidal rings, R/a = φ) - Surface standing wave (cyan, electron boundary layer) - Available valence sites (white glowing spheres) - Phase locking lines (gold) - Leptonic harmonics (electron, muon, tau) - Multi-angle views: Front, Top, Exploded, Leptonic Harmonics 3. GEOMETRY OF MATTER (PDF) - Complete geometric theory of the periodic table - Stable vortex packings: 1, 2, 6, 8, 10, 12, 14, 16, 18, 20, 24, 28, 32, 36, 50, 54, 82, 86, 126 - Gaps explained as geometric impossibilities - Valence as geometric availability 4. NUCLEAR STABILITY SIMULATION (Code + Thesis PDF) - Nuclear stability as phase equilibrium, not force equilibrium - Φ_critical = -51.1 mV as the universal threshold of matter stability - Fission as phase decoherence, fusion as phase locking - The electron as a surface pressure wave KEY FINDINGS:- The strong force is not a force — it is phase pressure- Neutrons are pressure regulators, not "glue"- Fission is phase decoherence, not nuclear division- Fusion is phase locking, not nuclear collision- Φ_critical = -51.1 mV is the universal threshold of matter- The electron is a surface pressure wave, not a point particle VALIDATION:- 100% proof obligations satisfied by GNATprove- Matches CODATA (proton mass, electron mass, fine-structure constant)- 118 elements simulated, stability predicted correctly- 200 materials reclassified by phase IMPORTANT NOTE ON THE 118 VISUAL PLATES:The 118 visual plates included in this deposit were generated using AI (Gemini)under the supervision of the author. They are illustrative representations ofthe V3 Architecture's geometric interpretation of the periodic table. They arenot experimental data nor computational simulations. Any visual inconsistencies or inaccuracies in the plates reflect the currentcapabilities of AI image generation and do not affect the underlying physicalcalculations. For exact numerical results, stability predictions, and formalproofs, please refer to the Ada SPARK source code, which is formally verifiedby GNATprove at 100% proof obligations. REFERENCE:Ψ_V3 = 48016.8 kg·m⁻² (Zenodo DOI: 10.5281/zenodo.20580979) LICENSE: LPV3 (License for the Protection of V3)COMMERCIAL USE: Requires explicit written permission from the author. AUTHOR:Dr. Benhadid Outail (ORCID: 0009-0003-3057-9543)Independent Researcher, Blida, AlgeriaEmail: mediconsulte@gmail.com
Predicción de la Eficiencia en Teleportación Cuántica en Presencia de Ruido Mediante Aprendizaje Automático e Inteligencia Artificial Explicable
Este repositorio recoge el código desarrollado para mi Trabajo de Fin de Grado (TFG) en Física en la Universidad Europea de Madrid. Este código da soporte a los experimentos y resultados presentados en el mismo. El código está desarrollado en el entorno de Google Colab y utiliza Qiskit y técnicas de Machine Learning para predecir la fidelidad de un estado teleportado en entornos ruidosos, empleando IA Explicable (XAI) para analizar qué tipos de ruido afectan más al protocolo. -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- This repository contains the code developed for my Bachelor's Thesis (TFG) in Physics at the Universidad Europea de Madrid. This code supports the experiments and results presented therein. The code is developed in the Google Colab environment and utilizes Qiskit and Machine Learning techniques to predict the fidelity of a teleported state in noisy environments, employing Explainable AI (XAI) to analyze which types of noise most significantly affect the protocol.
SMILE — Sustainable Methodology for Impact Lifecycle Enablement (v6.4.5): Operational Grammar for Auditable, Shared and Forkable Reality
Version 6.4.5 makes the one-page methodology overview readable. In v6.4.4 the figure was 1999 pixels compressed into a 6.6-inch text column — about 303 dpi, roughly 4-point type — and was attached separately, so the record opened on the image rather than the paper. Here the figure is given its own landscape page at full width, and six panels cut from it are repeated inside the sections they explain: the scale axis in 2.4, the AI Journey in 4.4, the Spin Twin Lifecycle in 4.7, the Outcome-to-Data chain in 5.2, the Alignment row in 5.5, and the Cycles and Use Cases rows in 6.1. Each panel lands between 75 and 154 dpi, so its smallest labels are two to four times larger than before. Banner sentences that span the full width of the figure are quoted verbatim in the panel captions rather than cut across panels. The overview is no longer a separate attachment; it is part of the text. No claim, reference or figure other than the overview is changed from v6.4.4. Version 6.4.4 restored the one-page methodology overview to the front of the paper, immediately after the abstract. The overview predates this specification: it was drawn approximately four years earlier in work with manufacturing industry, and it is the artefact from which the phase sequence, the cycles and the parallel AI journey in this paper descend. It is reproduced unchanged and is therefore not fully industry-agnostic; where its wording differs from the specification, the specification governs. SMILE (Sustainable Methodology for Impact Lifecycle Enablement) is a lifecycle methodology and design-science specification for constructing continuously learning digital twins as a consequence of impact-first project work. Version 6.4.x adds the operational grammar for auditable, shared and forkable reality: the Reality Branch Lifecycle ("no simulation becomes history"), the Reality Consensus Protocol, a canonical metamodel mapped to PROV-O, SOSA/SSN, OWL-Time, GeoSPARQL, ODRL, SHACL, NGSI-LD, ISO 23247, OpenUSD, MCP and A2A, an applicability and proportionality instrument, a spatial business and operating model, virtual-world rights and governance, a measurement framework, and fourteen falsifiable research propositions. The paper reframes Tuckman's group-development stages as concurrent capabilities on a shared reality substrate (grounding, situated contestation, protocolisation, closed-loop action, reconstitution), introduces Recorded Reality, the Epistemic Horizon, Reflexive Zoom, Reality-Defined systems, SEEN (Situation, Evidence, Effects, Needs) and the digital twin as a constitutional obligatory passage point. Version 6.4.3 completes an exhaustive verification of all 153 references, redraws the metamodel and proportionality figures, and adds an archival deployment record (2022) as lineage evidence. Note: the acronym expansion changed from "Interoperable" (v5.x) to "Impact" (v6.x); see the version history in the manuscript.
The data management crisis behind AI-based brain MRI diagnosis:Heterogeneity, governance, and reproducibility
AI-based brain magnetic resonance imaging (MRI) classifiers face a persistent gap between laboratory performance and clinical deployability. We argue the bottleneck is not model architecture but the absence of principled data management infrastructure. Drawing on our ongoing work in explainable multi-disease MRI classification, we identify five critical data management challenges: heterogeneity, governance misalignment, pipeline irreproducibility, annotation provenance, and demographic bias. To address these issues, we propose a minimal proof-of-concept metadata harmonization toolkit as a practical and tractable starting point for standardization. We further outline a validation protocol that measures field-level mapping coverage, clinician-reviewed preservation of phenotype and severity-score semantics, provenance completeness, and downstream classifier transfer across ADNI and OASIS-3. We invite the data management community to co-design the schema-mapping and query-optimization layers.
Vision-Based Machine Learning voor Chemische Procescontrole: Een Haalbaarheidsstudie over Meerdere Toepassingen
This research is motivated by the increasing complexity of chemical processes and the growing demand for robust, non-invasive monitoring and control solutions in the context of the ongoing digitalization of the chemical industry. Many critical aspects of chemical processes, such as phase behavior, flow patterns, and fouling, are inherently visual and therefore difficult to capture using traditional point measurements. By leveraging computer vision as a sensing modality, this thesis addresses a critical gap between observable process behavior and advanced, data-driven process control. The proposed methodology aims to support the transition towards more autonomous and intelligent operation by emphasizing robustness, interpretability, and seamless integration with existing industrial control frameworks, thereby facilitating the practical adoption of vision-based process control in the chemical sector. The research is structured around three primary objectives. The first objective focuses on the development of machine vision methods for classification and segmentation in chemical production environments characterized by continuous operation and high intrinsic safety standards. As a result, truly abnormal or failure-related events occur rarely, leading to highly imbalanced and limited datasets that pose significant challenges for conventional supervised learning approaches. Moreover, the creation of labeled datasets in the chemical sector is often prohibitively expensive, as it requires the involvement of domain experts to accurately interpret and annotate complex process phenomena. To address these constraints, two novel methods are proposed. The first method employs generative adversarial networks for anomaly detection, incorporating tailored cost functions and the structural similarity index to enable automated segmentation. This approach outperforms conventional supervised segmentation models trained for task-specific detection problems. The second method combines original and synthetically generated data to optimize classifier performance while quantitatively assessing generalization through an explainable AI framework. This strategy demonstrates superior performance compared to standard data augmentation techniques, increasing classification accuracy for the chemical foam classification task from 57% to 91%. The second objective focuses on developing a vision-based closed-loop control strategy for process management in the chemical sector. A laboratory setup simulating a chemical foaming production process, in which foam is continuously generated and a vision-based dosing system has been implemented to control foam volume, was established to develop a strategy capable of maintaining effectiveness even when precise control input accuracy cannot be guaranteed. The results demonstrate the feasibility of vision-based process control for automated anti-foaming agent dosing. Furthermore, a sensitivity analysis was conducted to evaluate the impact of the detection system's performance on the control solution. The analysis revealed that the precision of the detector has a limited effect on the system's ability to mitigate foam formation, whereas recall plays a more critical role: if recall dropped below 50%, the system was no longer able to effectively combat foam formation. The thesis concludes with the introduction of a generic framework for the development and deployment of machine vision-based control applications in chemical settings. This framework provides guidance on addressing data acquisition challenges, selecting suitable models, and integrating them into live production environments for automated decision-making. The framework was validated through four use cases at BASF Antwerpen, showcasing its applicability in real-world environments and gave way for several cost savings. The contributions of this thesis significantly advanced the field of vision-based process control, particularly within BASF Antwerpen. The findings emphasize the importance of robust model development, effective use of synthetic data, and the integration of machine vision systems into closed-loop control processes. These advancements offer tangible benefits for automation, efficiency, and cost reduction in chemical production environments and has been applied in four different use cases at BASF Antwerpen.
The Intent Bottleneck: When AI Compresses Delivery, Strategy Becomes the Constraint
This working paper argues that AI-native delivery compresses only the final stage of a six-stage chain running from an institution's covenant with those it serves to running software, so that by Amdahl's law the constraint moves upstream into how organisations form, prioritise, and preserve intent. It introduces intent provenance — the unbroken, traceable, durable, machine-checkable, and contestable chain from a stated obligation to the test that proves it in production — and the service covenant as the layer above vision and mission at the top of that chain. In closing it defines decision custody: an institution's demonstrated ability to reproduce, explain, defend, and answer for every consequential decision made in its name, human or machine, for as long as the consequences of that decision last. Decision custody is positioned as the institutional counterpart to decision sovereignty, a state-level concept defined elsewhere in the literature. The covenant is the promise; custody is the proof. Part of the Sovereign Digital Resilience series.
Opening the Black Box a Crack: A Interdisciplinary Mapping Review of Explainability and Interpretability of Black-Box Models
This article presents a narrative review of Explainability and Interpretability of Black-Box Models in the context of Artificial Intelligence. The literature on this topic has expanded substantially over recent decades, yet it remains fragmented across subfields, methods, and national research traditions. Drawing on an interpretive synthesis of representative contributions, the review reconstructs the historical development of the area, examines the conceptual foundations and definitional disputes that organize its debates, and maps the contemporary landscape of research, including the methodological shift toward data-intensive approaches and the institutional pressures that shape publication practice. Particular attention is given to the role of explainable AI and interpretability as organizing themes, and to the conditions under which findings from different research traditions can be brought into productive comparison. The review identifies three synthetic conclusions: the literature is cumulatively strong but organizationally weak; methodological pluralism is better understood as a resource than as a defect; and the growing practical salience of the topic raises the stakes of its unresolved conceptual questions. An agenda for future work is proposed, emphasizing integrative research designs, transparent synthesis practices, and the protection of definitional and infrastructural work on which cumulative progress depends. The article is intended as both a reference map for newcomers and a provocation for specialists in Artificial Intelligence.
AI-POWERED PREDICTIVE ANALYTICS FOR INTELLIGENT DECISION SUPPORT SYSTEMS
The digital landscape is undergoing a seismic shift, moving beyond the era of simple data storage into an age where the true value of information lies in its ability to forecast the future. AI-Powered Predictive Analytics for Intelligent Decision Support Systems is designed as a comprehensive guide to navigating this transition, blending the technical rigor of machine learning with the pragmatic needs of modern industrial and academic decision-making. In an increasingly complex world, the scope of global systems—spanning healthcare, finance, and manufacturing—has begun to exceed the limits of unassisted human intuition. This book explores the vital synergy between Artificial Intelligence (AI) and Decision Support Systems (DSS), illustrating how predictive modeling transforms raw, historical data into a strategic asset that anticipates trends, mitigates risks, and optimizes outcomes in real-time. While many existing texts focus solely on the mathematical algorithms of machine learning or the administrative management of information systems, this work bridges the gap by covering the entire system lifecycle. It takes the reader on a structured journey from the foundational theories of predictive analytics to the cutting edge of autonomous decision systems. Throughout these chapters, we delve into the intricate nuances of feature engineering, data governance, and the deployment of models within modern MLOps frameworks. Furthermore, the book provides deep technical explorations of supervised learning, ensemble methods, and deep learning architectures like CNNs and LSTMs, while placing a heavy emphasis on Explainable AI (XAI) to ensure that automated decisions remain transparent and trustworthy. The theory presented in these pages is anchored by extensive case studies that reflect both global and regional perspectives, providing a balanced view of how these technologies are implemented across diverse economic and regulatory environments. As we move toward a future of fully autonomous systems, the book addresses the critical challenges of algorithmic bias, data privacy, and the indispensable role of human-AI collaboration. This text is intended for researchers, data scientists, and business leaders alike—anyone who seeks to understand the strategic implications and technical requirements of building systems that are not only intelligent and efficient but also ethical and transparent. It is our hope that this roadmap serves as a vital resource for those looking to harness the analytical power of machines to solve the most pressing challenges of our time.
From tech blogs
See all →Introducing WeatherNext 3, our most advanced and accurate global weather AI model
From MIT to IBM, expediting AI and quantum deployment
MIT affiliates engage with the MIT-IBM Computing Research Lab to bring rigorous theory to production systems.
System helps humans predict when self-driving cars will make mistakes
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
Planetary prediction engine: Automating global models via Earth AI
Earth AI
Piloting the world's first double-blind AI evaluations
Piloting the world's first double-blind AI evaluations