Skip to content
Open access

Explainable image aesthetic assessment via semantic segmentation-guided cross-attention fusion

Sep 2026 · Scientific Reports · Vol 16 · 0 citations · 34 references

Abstract

Most existing image aesthetic assessment methods rely on global representations or simple aggregation of local features. This makes it difficult to model semantic region information and provide region-level explanations. To address this issue, we propose a Semantic Segmentation-Guided Cross-Attention Fusion (SSG-CAF) method for explainable image aesthetic assessment. The method first uses SegFormer-B0 to generate semantic masks for eight semantic categories, thereby introducing explicit semantic and spatial priors. It then combines region-level features with global visual context. A cross-attention mechanism is used to model dependencies between semantic regions and the global representation. This design supports the representation of high-level aesthetic cues, including compositional balance, subject prominence and regional harmony. SHAP is further used for region-level attribution analysis. Experiments on the cleaned AVA test set showed that SSG-CAF achieved an MSE of 0.3412 ± 0.0150, an MAE of 0.4565 ± 0.0103, a PLCC of 0.6517 ± 0.0029 and an SRCC of 0.6370 ± 0.0034. Compared with the selected baselines, SSG-CAF showed stronger correlations with human aesthetic ratings under the cleaned AVA protocol. It also produced post-hoc, quantitative region-level attribution evidence that describes model behaviour.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.