Deep fake detection in face images and video by tuning a pretrained CLIP model
Abstract
This study develops a video deepfake detection system that addresses the critical challenge of cross-dataset generalization in real-world scenarios. It adopts CLIP (Contrastive Language-Image Pre-training) as the foundation, leveraging its strong vision-language representations for detecting subtle facial manipulations. A parameter-efficient fine-tuning approach is proposed that achieves superior cross-dataset performance while updating only a minimal fraction of model parameters. The CLIP ViT-B model, utilizing the LayerNorm tuning, achieves an average video-level AUC of 90.9% across seven benchmark datasets, while training only 0.0344% of the total parameters. This approach is also extended to CLIP-STAN, a temporal video detector, achieving competitive results with 9.46% trainable parameters. The work integrates CLIP with LayerNorm tuning into the Deepfake Bench framework, enabling comprehensive cross-dataset generalization analysis. Notably, the implemented solution - base CLIP ViT-B with LayerNorm tuning - matches more complex CLIP variants in generalization performance, underscoring the significance of high-quality pre-training over architectural complexity.