Skip to content
Preprint

AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

Aug 2026 · 0 citations · 90 references
Computer Science

TL;DR

AlignFace is proposed, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling and significantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics.

Abstract

Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While perceptual evaluation has progressed from signal-based heuristics to representation-based metrics, current approaches are limited to behavioral modeling without cognitive alignment. They rely on implicit and spurious relations while assuming a universal observer, failing to account for inherent variations across diverse human populations. This leads to inaccurate evaluative models of stakeholders and misleading guidance for generative model debugging. Rather than treating perception as a black box, we leverage scientific findings from cognitive psychology of human face similarity perception: dependence on facial featural and configural attributes, nonlinear psychophysical response scaling, and own-group biases. We introduce the FACETS dataset and propose AlignFace, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling. It employs visual-language modeling (VLM) to encode paired face images and text-based attributes, gated cross-attention (CA) to extract attribute-specific facial difference representations, concept bottleneck modeling (CBM) to constrain reasoning via interpretable face attributes, and neural generalized additive model (GAM) to model their nonlinear influence. Experiments found AlignFace significantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics. By bridging learned representations and human cognitive processes, this work enables more transparent and aligned perceptual evaluation metrics for face images.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. Ho...

Laurent Colbois, Sébastien Marcel · 0 citations
Preprint Sep 2026

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

Judging image quality is not only ecologically relevant to everyday human tasks, but also underpins many machine vision tasks such as image generation. This paper proposes a framework to understand the inherent perceptual space underlying image quality judgment in humans. We propose a multi-dimensional observer model t...

Sheng Zhao, Wei-Kai Lin, Yu-Hao Zhu · 0 citations
Preprint Aug 2026

More Accurate, Less Human: Gestalt Grouping in Vision Models

A behavioral battery is introduced that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out, revealing aspects of perceptual organization that conventional performance metrics fail to disting...

Sudhanva Manjunath Athreya, S. Malladi · 0 citations
Open access

Integrating Domain Knowledge for Robust and Interpretable Deep Neural Networks in Facial Affect Recognition

Facial affect recognition is a key component of human-centered AI, enabling systems to respond appropriately to human nonverbal signals. The deployment of deep learning for facial affect recognition is critically hindered by vulnerability to data-driven biases and a lack of transparency. This thesis addresses these cha...

Ines Rieger · 0 citations
Preprint Aug 2026

Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models

This work uses simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models, exposing face embeddings as semantically and visually rich biometric representations for web-scale foundation models.

Fizza Rubab, Yi-Ying Tong, Arun Ross · 0 citations
Sep 2026

Impact of Facial Variability on Shape-Only Landmark-Based Face Verification

Facial landmarks are commonly used in face recognition for alignment, pose normalization, and quality assessment, but they can also serve as an interpretable representation of facial shape. This paper studies shape-only 1:1 face verification using dense 2D facial landmarks, without texture descriptors, deep face embedd...

Emma Machacova, P. Helebrandt · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.