Skip to content
Preprint

Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors

Jul 2026 · 0 citations · 45 references
Computer Science

Abstract

Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only consider predicting a binary label for whether the input is real or fake, lacking the justification required for real-world applications like legal proceedings. Explainable DF Detection has emerged to address this limitation, but existing techniques frequently fall short by either relying on human annotations for precise artifact localization or generating superficially plausible textual explanations without grounding. This work investigates the use of post-hoc explainable AI (XAI) to analyze the decision-making process of state-of-the-art black-box DF detectors. Specifically, we employ Encoding-Decoding Direction Pairs (EDDP), a technique suitable for uncovering the concept space of DF detectors (their semantic vocabulary) as well as the mechanism for writing and reading concept information to and from internal representations. Our analysis reveals previously hidden real and fake features learned implicitly during detector training, offering nuanced explanations unattainable through conventional methods. This enables global model understanding, spatially aware concept localization, and counterfactual what-if analysis, all contributing to a deeper comprehension of DF detection strategies.

View source

Similar papers

Preprint Jul 2026

Explainable Deepfake Detection Challenge

The Explainable Deepfake Detection Challenge at ACM Multimedia 2026 is designed to benchmark this joint capability of classification metrics with semantic similarity, simplicity, and intent-aware grounding metrics that assess whether explanations identify the relevant manipulated entities and supporting visual evidence.

Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh et al. · 1 citation
Preprint Aug 2026

Latent Fact-Checking: Detecting Misinformation through Activation Engineering

Findings provide evidence that truthfulness is a structured, linearly separable concept in the latent space of pretrained language models, and point toward interpretability-driven misinformation detection as a practical complement to retrieval-based pipelines.

P. Barcelos, Otávio Parraga, Marcelo M. Mussi et al. · 0 citations
Preprint Aug 2026

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.

Chidaksh Ravuru, Shashank Srivastava · 0 citations
Review Jul 2026

Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art

A survey and comparative analysis of NLP-based Automatic Deception Detection focusing on the legal domain and the evolution from feature-based machine learning to Large Language Model (LLM) approaches are presented, showing strong domain sensitivity.

T. Samaradiwakara, Nisansa de Silva, George C. Lobb · 0 citations
Preprint Jul 2026

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

This work introduces XPlainVerse, a large-scale benchmark designed for joint deepfake detection and human-centered explanation, and proposes novel metrics, EntityScore and EvidenceScore, that measure reasoning fidelity by checking whether explanations correctly identify manipulated entities and visual evidence.

Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh et al. · 0 citations
Jul 2026

Explainable Prompt Injection Detection using Sentence Embeddings, Random Forest, and Word-Level Attribution

A explainable framework that detects the attacks using a Random Forest (RF) classifier along with semantic sentence embedding and SHAP (Shapley Additive Explanations) is used together with a wordlevel attribution mechanism to give human-legible explanations for the model predictions, and to emphasize harmful parts within the prompt.

M. Ahmad, M. Asif, Aoun E. Muhammad et al. · 0 citations