Skip to content
Preprint

Foundation and Multimodal Large Language Models for Face Presentation and Morph Attack Detection

Aug 2026 · 0 citations · 108 references
Computer Science

TL;DR

The experiments show that FMs and MLLMs can achieve significant performance for PAD and MAD and the fine-tuned models achieve state-of-the-art detection performance in cross-dataset evaluation, indicating that general-purpose pretrained representations carry substantial attack-relevant information.

Abstract

Face recognition systems are increasingly deployed in security-critical applications, yet they remain vulnerable to presentation and morph attacks. Presentation attack detection (PAD) and morphing attack detection (MAD) are therefore essential components of trustworthy face biometrics. Despite advancements in PAD and MAD methods, existing detectors suffer from limited generalization and degrade in cross-dataset evaluation. In this paper, we systematically investigate whether general-purpose foundation models (FMs) and multimodal large language models (MLLMs) encode PAD-relevant and MAD-relevant information, and how such models can best be deployed for both tasks. We study five approaches with increasing access to the internal information of the model: (i) zero-shot prompting of off-the-shelf MLLMs; (ii) training a shallow model on the next-token logit probabilities at the output of the MLLM; (iii) parameter-efficient fine-tuning on task-specific question-answer data, yielding two specialized MLLMs, called PADLLM and MADLLM, which additionally provide textual reasoning for their decisions; (iv) linear probing of frozen vision encoders; and (v) fine-tuning of vision encoders of FMs and MLLMs. We benchmark 16 open-weight MLLMs and 30 vision encoder backbones on four PAD datasets (MSU-MFSD, CASIA-FASD, Replay-Attack, and OULU-NPU) and four MAD datasets (FFHQ, FRGC, FRLL, and FERET). Our experiments show that FMs and MLLMs can achieve significant performance for PAD and MAD. In addition, the fine-tuned models achieve state-of-the-art detection performance in cross-dataset evaluation, indicating that general-purpose pretrained representations carry substantial attack-relevant information. Source code of all our experiments will be publicly released.

View source

Similar papers

Preprint Aug 2026

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection

Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-dataset evaluation. Variations in sensors or lighting conditions can reduce the effectiveness of detector...

P. Lorenz, Anjith George, Marcel Sébastien · 0 citations
Preprint Aug 2026

MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection

A few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems, and validating the robustness of semantic anchoring compared to models trained on massive short-text data.

Xiao-Yong Yu, Rong-Zhen Li, Shu-Ming Shi et al. · 0 citations
Sep 2026

Design and Implementation of Hybrid CNN-Transformer Architecture for Face Anti-Spoofing

Biometric face recognition systems are progressively being used in security-sensitive areas and are vulnerable to presentation attacks (PAs) such as printed photograph recaptures, video replay attacks and 3D silicone mask attacks. Face Anti-Spoofing (FAS) also known as liveness detection is the first line of defence ag...

Shivakumar Dalali · 0 citations
Preprint Aug 2026

Adversarial Attacks on Deep OCR Systems

Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black...

Wenbo Sun, Hong-Zong Li, Yanyun Wang et al. · 0 citations
Conference Open access Aug 2026

XSA-Mad: Cross-Modal Semantic Alignment for Morphing Attack Detection

XSA-MAD, a CLIP-based multimodal framework that explicitly models semantic inconsistencies between bona-fide and morphed faces, is proposed, a CLIP-based multimodal framework that consistently outperforms existing methods under high-fidelity generative attacks.

Jie Jin, Mahiro Tokumasu, Yushi Makino et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.