Skip to content
Review Open access

A systematic review of attention duality in deep visual models for performance and faithfulness tradeoffs

Sep 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 187 references

Abstract

While deep learning models have achieved remarkable performance in computer vision, their internal decision logic often remains opaque. Attention mechanisms serve a dual function: dynamically prioritizing discriminative features to maximize predictive performance while simultaneously generating visual heatmaps for explanation. However, the extent to which attention maps faithfully reflect true computational reasoning remains controversial. This study presents a Systematic Literature Review (SLR) of 168 primary studies evaluating attention-based explanation faithfulness across Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). We propose a structured taxonomy categorizing the literature into four distinct paradigms defined along two orthogonal axes: architectural integration role and interpretability design intent (spanning from unconstrained implicit byproducts to explicit attribution and causal verification). Quantitative synthesis reveals that 57.14% of primary studies leave the performance–faithfulness relationship unexamined, 41.07% report architectural synergy, and merely 1.79% document an explicit trade-off—highlighting significant historical publication bias. Furthermore, this review traces an evolutionary shift in evaluation methodologies from qualitative visual inspection toward objective human-grounded alignment and perturbation-based causal benchmarks. By addressing key methodological gaps, we establish a standardized reporting checklist and an actionable roadmap to engineer deep visual models that achieve verifiable alignment between predictive accuracy and explanation faithfulness.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.