Skip to content
Preprint

SCOUT: Semantic Concept Discovery for Open-Vocabulary Editing of face Recognition Templates

Aug 2026 · 0 citations · 45 references
Computer Science

TL;DR

Experiments with face recognition models using CNN, ViT, and Swin backbones show that SCOUT discovers interpretable concepts beyond standard attribute labels and enables controllable, identity-aware template manipulation with negligible impact on identity matching.

Abstract

Face recognition templates are compact identity representations, yet they also encode rich semantic information about facial appearance. Prior work has shown that templates can be inverted to images or indirectly manipulated through image-editing pipelines, but direct semantic editing in template space remains largely unexplored. Existing interpretability methods for face recognition often rely on manual neuron inspection or predefined attribute labels, limiting scalability and semantic flexibility. To address this gap, we propose SCOUT (Semantic Concept Discovery for Open-VocabUlary Editing of Face Recognition Templates), an end-to-end framework for discovering and directly manipulating semantic concepts in face recognition templates using mechanistic interpretability. SCOUT learns sparse template representations, generates semantic hypotheses for latent features from natural-language descriptions, and validates their stability. The resulting features act as controllable semantic directions for direct editing, avoiding costly edit--re-encode pipelines. Experiments with face recognition models using CNN, ViT, and Swin backbones show that SCOUT discovers interpretable concepts beyond standard attribute labels and enables controllable, identity-aware template manipulation with negligible impact on identity matching. We further show that edited templates can subsequently be decoded with independent inversion models for visualization and evaluation.

View source

Similar papers

Preprint Aug 2026

Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models

This work uses simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models, exposing face embeddings as semantically and visually rich biometric representations for web-scale foundation models.

Fizza Rubab, Yi-Ying Tong, Arun Ross · 0 citations
Preprint Aug 2026

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, agai...

Guray Ozgur, Mustafa Efe Tamyapar, N. Damer et al. · 0 citations
Conference Open access Sep 2026

Open-Vocabulary Object 6D Pose Estimation via Modulated Textual Semantics

Estimating the 6D pose of novel objects without CAD models or video sequences remains a challenging problem. Recent works explore text-driven approaches to address this challenge in an open-vocabulary manner. However, these methods typically treat text embeddings as static priors, which lack the flexibility to adapt to...

Zixuan Sun, H. Shuai, Qing-Shan Liu · 0 citations
Open access Aug 2026

Compact Occlusion-Robust Facial Expression Recognition via Clean-Anchored Hard Occlusion Fine-Tuning

Control comparisons and ablations indicate that the retained model has the most favorable observed cleanness–robustness trade-off among the tested epoch-matched alternatives; however, fixed-checkpoint comparisons on Occlusion-RAF-DB are not significant after Holm correction, while broader cross-domain validation remain...

Xue-Feng Zhao, Yi-Xuan Dong, Zhao-Man Zhong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. Ho...

Laurent Colbois, Sébastien Marcel · 0 citations
#machine learning Preprint Sep 2026

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such a...

Shuang Liang, Le-Jun Liao, Shi-Yuan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.