• English
  • Deutsch
  • Log In
    Password Login
    Research Outputs
    Fundings & Projects
    Researchers
    Institutes
    Statistics
Repository logo
Fraunhofer-Gesellschaft
  1. Home
  2. Fraunhofer-Gesellschaft
  3. Scopus
  4. Vision Transformers for Face Recognition Need More Registers
 
  • Details
  • Full
Options
2026
Conference Paper
Title

Vision Transformers for Face Recognition Need More Registers

Abstract
Recent advances in Vision Transformers (ViTs) for face recognition (FR) have moved beyond the standard CLS-token paradigm. In this paradigm, a special classification token (CLS) is prepended to the patch embeddings and used as a representation of the input for downstream tasks. An alternative approach, Concatenated Patch Embeddings (CPE), instead leverages all patch tokens by concatenating them into a single vector, which is then projected into a compact face representation. CPE has been shown to improve recognition performance in comparison to CLS-based ones, but our qualitative analysis of attention maps showed the presence of artifacts that limit their interpretability. To address this issue, we incorporate register tokens, learnable tokens concatenated to the initial patch embeddings, and processed jointly through the ViT encoder blocks. This mechanism has been shown to produce more structured and interpretable attention maps compared to baseline ViT. We empirically demonstrate that these artifacts consistently appear across various ViT backbones, including small and large models, and that introducing register tokens effectively mitigates them. Adding four or eight registers significantly enhances interpretability, with eight registers providing the highest verification accuracies and smoothest attention structures. Our resulting model, ViT-8R, corresponds to a CPE-based ViT-B architecture augmented with eight register tokens achieves state-of-the-art performance among ViT-based FR models on large-scale IJB-B and IJBC benchmarks. Also, ViT-8R produces substantially clearer attention maps compared with the baseline model, which offer deeper insight into the model's attention behavior (https://github.com/TaharChettaoui/ViT-FR-Registers).
Author(s)
Chettaoui, Tahar
Fraunhofer-Institut für Graphische Datenverarbeitung IGD  
Ozgur, Guray
Fraunhofer-Institut für Graphische Datenverarbeitung IGD  
Loureiro Caldeira, Maria Eduarda
Fraunhofer-Institut für Graphische Datenverarbeitung IGD  
Damer, Naser  
Fraunhofer-Institut für Graphische Datenverarbeitung IGD  
Boutros, Fadi  orcid-logo
Fraunhofer-Institut für Graphische Datenverarbeitung IGD  
Mainwork
IEEE 20th International Conference on Automatic Face and Gesture Recognition, FG 2026  
Funder
Bundesministerium für Forschung, Technologie und Raumfahrt  
Conference
International Conference on Automatic Face and Gesture Recognition 2026  
DOI
10.1109/FG67764.2026.11556970
Language
English
Fraunhofer-Institut für Graphische Datenverarbeitung IGD  
  • Cookie settings
  • Imprint
  • Privacy policy
  • Api
  • Contact
© 2024