Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.
Abstract
The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.
This paper proposes an end-to-end Transformer-based framework, termed Progressively Explicit Query Network (PEQNet), for multi-face forgery detection and localization, and introduces triple contrastive learning to model the mutual exclusivity among real, fake, and background regions.
Peng-Wen Dai, Xiaomeng Wen, Feiyang He et al.· ACM Transactions on Multimed...· 0 citations
Generalizable face forgery detection has become a critical problem in multimedia forensics as modern face manipulation techniques can generate increasingly realistic facial content. Vision foundation models provide a promising basis for this problem, but existing detectors usually rely on the last-layer visual feature,...
Yi-Meng Zhao, Shuo Zhu, Jia-Lang Liu et al.· 2026 12th International Conf...· 0 citations
A few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems, and validating the robustness of semantic anchoring compared to models trained on massive short-text data.
Xiao-Yong Yu, Rong-Zhen Li, Shu-Ming Shi et al.· 0 citations
The proposed Optimized Multi-level Mixed Attention-enabled Hybrid Learning-based Bidirectional Gradient (OM2AHL-BiG) model results in enhanced detection, adaptability to various forgery types, and improved interpretability, making it an effective solution for detecting deepfake and intra-frame video forgeries.
A novel framework, UVIF, that utilizes additional annotated images to provide fine-grained supervision for detecting partial forgeries in videos, which outperforms state-of-theart methods in detecting partially forged videos while introducing no additional computational overhead is proposed.
Haotian Liu, Y. Liu, Guoying Zhao et al.· 0 citations
Deepfake technology, powered by deep learning models, enables the synthesis of highly realistic facial images and videos. However, in recent years, the misuse of deepfakes has posed severe challenges to both individual privacy and social trust. Consequently, this paper systematically reviews research pertaining to deep...
Yu-Jing Zhou· ITM Web of Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.