Skip to content
Preprint

MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection

Aug 2026 · 0 citations · 52 references
Computer Science

TL;DR

A few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems, and validating the robustness of semantic anchoring compared to models trained on massive short-text data.

Abstract

Facial biometric recognition systems currently face compound threats intertwining generative AI and high-fidelity physical spoofing. Existing defenses suffer from systemic bottlenecks, including poor generalization, non-auditable reasoning, and reliance on massive, low-quality datasets. To address these challenges, we propose Multimodal Large Language Models (MFAD) for face anti-spoofing detection, an explainable reasoning system for Unified Face Anti-Spoofing Detection (UFAD), accompanied by a semantic-level annotation benchmark. Unlike methods relying on external tools or coarse alignment, MFAD activates the intrinsic reasoning capabilities of Multimodal Large Language Models (MLLMs) via a fine-grained pixel-semantic anchoring mechanism. This eliminates localization hallucinations and ensures auditable reasoning paths. We introduce a cross-attack semantic-level unified annotation paradigm: by annotating only 1,000 precise masks per attack category, we generate reasoning evidence chains strictly corresponding to spoofed regions. Supervised fine-tuning on the Qwen-VL foundation model demonstrates that, using limited high-quality samples, the system achieves a 40-50% relative reduction in in-domain ACER and restricts cross-domain performance degradation to within 11.62%/5.23%, significantly outperforming existing frameworks. Furthermore, under white-box adversarial attacks, detection accuracy drops by only 3.2%, validating the robustness of semantic anchoring compared to models trained on massive short-text data. Domain practitioners rated the evidence reliability of reasoning paths at 4.57/5, with inference latency satisfying real-time deployment requirements. These results confirm that a few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems.

View source

Similar papers

Open access Sep 2026

Reliability-aware vision-language face anti-spoofing via progressive semantic reorganization

Face Anti-Spoofing (FAS) is a crucial task for securing face recognition systems, yet its cross-domain generalization remains challenging. Recently, vision-language methods built upon pretrained models such as CLIP have shown promising performance in addressing these cross-domain scenarios. Nevertheless, existing appro...

Xiao-Meng Wei, Wen-Zhong Yang, Ya-Bo Yin et al. · 0 citations
Open access 2026

Semantically Anchored Test-Time Domain Generalization for Face Anti-Spoofing

The Semantic-Anchored Test-Time Domain Generalization (SA-TTDG) framework is introduced, introducing a Text-Anchored Style Projection (TASP), which utilizes rich linguistic priors from Vision-Language Models (VLMs) to initialize and strongly constrain learnable style bases.

Xiaosong Chang, Liang Shi, Ao Zhang · 0 citations
Preprint Aug 2026

Foundation and Multimodal Large Language Models for Face Presentation and Morph Attack Detection

The experiments show that FMs and MLLMs can achieve significant performance for PAD and MAD and the fine-tuned models achieve state-of-the-art detection performance in cross-dataset evaluation, indicating that general-purpose pretrained representations carry substantial attack-relevant information.

H. O. Shahreza, A. H. Khan, P. Lorenz et al. · 0 citations
Conference Aug 2026

Intelligent multimodal face anti-spoofing detection system based on biometrics

A multi-dimensional feature fusion-based face anti-spoofing detection method based on the YOLOv8 architecture that integrates dynamic optical flow features with static texture analysis and achieves robust detection through the fusion of spatial and temporal cues.

Yanhua Liang, Pengcheng Zhou, Hongmei Qin et al. · 0 citations
Conference Open access Aug 2026

XSA-Mad: Cross-Modal Semantic Alignment for Morphing Attack Detection

XSA-MAD, a CLIP-based multimodal framework that explicitly models semantic inconsistencies between bona-fide and morphed faces, is proposed, a CLIP-based multimodal framework that consistently outperforms existing methods under high-fidelity generative attacks.

Jie Jin, Mahiro Tokumasu, Yushi Makino et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.