Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only consider predicting a binary label for whether the input is real or fake, lacking the justification required for real-world applications like legal proceedings. Explainable DF Detection has emerged to address this limitation, but existing techniques frequently fall short by either relying on human annotations for precise artifact localization or generating superficially plausible textual explanations without grounding. This work investigates the use of post-hoc explainable AI (XAI) to analyze the decision-making process of state-of-the-art black-box DF detectors. Specifically, we employ Encoding-Decoding Direction Pairs (EDDP), a technique suitable for uncovering the concept space of DF detectors (their semantic vocabulary) as well as the mechanism for writing and reading concept information to and from internal representations. Our analysis reveals previously hidden real and fake features learned implicitly during detector training, offering nuanced explanations unattainable through conventional methods. This enables global model understanding, spatially aware concept localization, and counterfactual what-if analysis, all contributing to a deeper comprehension of DF detection strategies.
The Explainable Deepfake Detection Challenge at ACM Multimedia 2026 is designed to benchmark this joint capability of classification metrics with semantic similarity, simplicity, and intent-aware grounding metrics that assess whether explanations identify the relevant manipulated entities and supporting visual evidence.
Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh et al.· 1 citation
Findings provide evidence that truthfulness is a structured, linearly separable concept in the latent space of pretrained language models, and point toward interpretability-driven misinformation detection as a practical complement to retrieval-based pipelines.
P. Barcelos, Otávio Parraga, Marcelo M. Mussi et al.· 0 citations
U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.
A survey and comparative analysis of NLP-based Automatic Deception Detection focusing on the legal domain and the evolution from feature-based machine learning to Large Language Model (LLM) approaches are presented, showing strong domain sensitivity.
T. Samaradiwakara, Nisansa de Silva, George C. Lobb· 0 citations
This work introduces XPlainVerse, a large-scale benchmark designed for joint deepfake detection and human-centered explanation, and proposes novel metrics, EntityScore and EvidenceScore, that measure reasoning fidelity by checking whether explanations correctly identify manipulated entities and visual evidence.
Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh et al.· 0 citations
A explainable framework that detects the attacks using a Random Forest (RF) classifier along with semantic sentence embedding and SHAP (Shapley Additive Explanations) is used together with a wordlevel attribution mechanism to give human-legible explanations for the model predictions, and to emphasize harmful parts within the prompt.
M. Ahmad, M. Asif, Aoun E. Muhammad et al.· ICACNC 2026 Proceedings· 0 citations